Work

IISc · MULTI-AGENT · 2026

Multi-Agent Code Review & Auto-Debugging System

A LangGraph multi-agent system that reviews a GitHub pull request with four agents in parallel (security, bugs, style and performance) and commits fixes back to the pull request. Built in the IISc Executive Education Program in Agentic & Generative AI.

Context
IISc Bengaluru · Executive Education Program – Agentic & Generative AI
My role
Architected the multi-agent platform; built the RAG-powered code intelligence system
Year
2026
Stack
LangGraph · FastAPI · ChromaDB · Claude · React

Context

I built this system in the Executive Education Program in Agentic & Generative AI at the Indian Institute of Science (IISc) Bengaluru, which I completed in 2026. The goal was to automate secure code review and AI-assisted debugging workflows.

My role

I architected the multi-agent platform and built the RAG-powered code intelligence system behind it.

How it works

A review starts from a pull-request URL or a GitHub webhook. The FastAPI service accepts it at once (202 Accepted), runs the review in the background and streams progress to a React dashboard.

  • Supervisor–Worker orchestration. A LangGraph orchestrator acts as the supervisor. It fetches the pull request's changed source files through the GitHub REST API, then fans the work out to four analysis agents that run in parallel: security, bug detection, style and performance.
  • Static checks first, then the LLM. Each agent pairs a cheap, deterministic check with an LLM (Claude, at temperature 0). The security agent retrieves the most relevant OWASP and CWE references from a ChromaDB vector store (RAG). The bug and performance agents run Python AST analyzers, and the style agent runs Ruff. Their results go to the LLM as hints, so the model spends its effort on what static checks can't see.
  • One findings contract. Every agent returns findings in the same JSON shape: severity, confidence, line range, suggestion and CWE ID. LangGraph reducers merge the parallel results into one review.
  • Patch generation. If any finding is medium severity or higher, a fix agent asks the LLM for corrected code, checks that the code still compiles, and commits the fixes to the pull request's own branch: one commit per category.
  • Test stage. The graph ends with a test stage that hands back the pull-request link. In the current build, generating tests there is switched off.
Multi-agent code review architectureA pull request URL or a GitHub webhook reaches a FastAPI service, which accepts the review at once and runs it in the background. A LangGraph orchestrator fetches the changed files through the GitHub API and fans out to four analysis agents that run in parallel: security, which retrieves OWASP and CWE references from a vector store; bug detection and performance, which combine Python AST checks with an LLM; and style, which combines Ruff with an LLM. Their findings are merged. If any are medium severity or higher, a fix agent generates patches, checks that they compile and commits them to the pull request branch through the GitHub API, one commit per category. Progress streams live to a React dashboard over server-sent events. LangFuse traces the run, it ships with Docker, and it was evaluated with RAGAS.Pull requestPR URL or GitHub webhookFastAPI serviceaccepts at once · runs in backgroundLangGraph orchestratorfetches changed files · GitHub APIfan-outANALYSIS AGENTS · IN PARALLELSecurityRAG · OWASP / CWEBug detectionAST checks + LLMStyleRuff + LLMPerformanceAST checks + LLMfan-inMerge findingsseverity · confidence · lines · CWEif medium or higherFix agentpatch · compile check · commit per categoryvia the GitHub APIPull request branchfixes ready for human reviewlive progress → React dashboard (SSE)LangFuse tracing · Docker · RAGAS evaluation
The orchestrator is the supervisor; the agents are its workers and run in parallel.

Key decisions and trade-offs

1. Parallel agents, merged with reducers

Decision. The four analysis agents run at the same time and write to shared state through LangGraph reducers.

Trade-off. A review takes about as long as its slowest agent, not the sum of all four: typically 30–60 seconds, depending on the size of the pull request and the model's latency. The cost is that concurrent writes need explicit merge rules.

2. Static analysis as hints for the LLM

Decision. The AST analyzers and Ruff run before the LLM, and their findings go into its prompt. Duplicate style findings are removed.

Trade-off. Deterministic checks are fast and never hallucinate, and they let the model focus on deeper issues. However, the AST analyzers cover Python only, so for other languages the agents rely on the LLM alone.

3. RAG that degrades gracefully

Decision. The security agent grounds its review in OWASP and CWE references retrieved from ChromaDB. If the vector store is unavailable, it falls back to a built-in set of core references.

Trade-off. The fallback is smaller, so findings can be less specific, but a review never fails because the vector store is down.

4. Commit through the GitHub API, one commit per category

Decision. There is no local clone. Each fix is committed with the Git Data API (blob, tree, commit, then a ref update) in a fixed order: security, bugs, style, performance.

Trade-off. Each category can be reviewed or reverted on its own, and there is no clone time or disk to manage. In return, a compile check is the only safety net before a commit. The scope is capped at medium-severity findings and above, and at 10 files per category. That is why the fixes land on the pull request for a person to review, never straight on the main branch.

5. In-process first, a queue later

Decision. The first version runs each review as a background task inside the API process. A Redis/ARQ worker is already designed for the next step.

Trade-off. The system is simple to run and debug, but a restart loses any review in flight until reviews move to the queue.

Reliability and security

  • GitHub webhooks are verified with HMAC-SHA256 signatures, and the API is rate-limited per client.
  • One failing file never stops an agent. Typed errors carry their own retry policy: for example, exponential back-off for GitHub and LLM errors, and no retry when a prompt is too long.
  • Correlation IDs run through structured logs, and LangFuse traces the agent and LLM calls.

Evaluation

I evaluated agent performance across four scenarios: code review, bug fixing, patch generation and test generation. The evaluation combined automated evaluation metrics, test execution and RAGAS-based evaluation.

Stack

Python · LangGraph · LangChain · Claude · ChromaDB (RAG over OWASP / CWE) · FastAPI · server-sent events · React · Redis / ARQ · LangFuse · Docker · RAGAS

Next case study: Expense Analyzer