Codebase Audit
Autonomous codebase health analysis powered by the trace graph. Get architecture reports, gap analysis, tech debt summaries, and actionable recommendations — all generated from your existing enrichment data.
Overview
AutoAudit analyzes your trace graph (nodes, edges, augmentations, epistemic enrichment, modules, and atlas) to produce structured findings and optional LLM-generated reports. It is a finalize pipeline stage (stage 14 of 15, running in finalize wave 2 after deep enrichment), so it runs automatically when the project's finalize group is set to auto — and you can also re-trigger it on demand via the dashboard Audit panel, the MCP prep_audit tool, or the REST stage-run endpoint.
AutoAudit V2 transforms SourcePrep from a "passive observer" into an "active taskmaster". Findings are categorized into flat tabs (Architecture, Quality, Coverage, Tech Debt), prioritized, and include concrete actionable items. You can select findings and click "Copy AI Command" to instantly hand off the context assembly to your AI via MCP.
6 findings total
1 critical · 2 warning · 2 info
Last run: 7/20/2026, 5:59:15 PM
Top Findings
By Module
Live preview: AutoAudit findings organized by severity and category with AI handoff.
Two Tiers
Pure graph queries. No LLM needed. Runs in <2 seconds. Produces structured findings (JSON).
LLM-generated markdown reports. Uses your configured large model. Produces 5 documents.
Quick Start
CLI
Audits are triggered via the MCP prep_audit tool, the REST API, or the dashboard's Audit panel — not a direct CLI command. To read findings from the most recent run from a terminal, use prep opportunities:
# List actionable findings from the latest audit
prep opportunities
# Filter by minimum priority
prep opportunities --priority P1
# Export as SARIF for CI ingestion
prep opportunities --format sarif
# Pipe a high-priority subset to an AI agent prompt
prep opportunities --priority P0 --format ai_promptWhat a PR-sanity audit looks like when an agent calls prep_audit mid-review.
MCP Tool
One MCP tool, prep_audit, exposes all audit functionality across Cursor, Windsurf, Claude Code, and any MCP-compatible editor. The behavior is selected via the action parameter:
| Action | What it does |
|---|---|
| scan (default) | Run analyzers and return structured findings. Set synthesize: true to also generate LLM reports. |
| report | Read a specific generated report by name (AUDIT_SUMMARY, ARCHITECTURE_ANALYSIS, GAP_ANALYSIS, COMPONENT_INVENTORY, TECH_DEBT_REPORT). |
| refactor | Context-assembly handoff. Pass finding_ids (e.g. ["ARCH-1", "QUAL-2"]) and get back the structural trace graph for all affected files, priming the AI for an immediate refactor. |
| verify | Re-run a named subset of analyzers to confirm recent edits actually resolved the finding. |
| antibodies | List immune-system defenses derived from concept assertions. |
Findings can also be passed in via the findings parameter to enrich external lint output (ruff, eslint, semgrep, SARIF) with structural context — see the Audit Enrichment guide.
REST API
| Method | Endpoint | Purpose |
|---|---|---|
| POST | /projects/{id}/pipeline/stages/audit/run | Trigger audit as a finalize stage through the orchestrator (runs Tier 1, then attempts Tier 2 synthesis; Tier 2 is skipped only on synthesis failure) |
| GET | /projects/{id}/audit/status | Check progress and last run metadata |
| GET | /projects/{id}/audit/findings | Raw structured findings (filterable) |
| GET | /projects/{id}/audit/reports | List generated report documents |
| GET | /projects/{id}/audit/report/{name} | Read a specific report |
Built-in Analyzers
AutoAudit ships with 11 analyzers. Each reads the trace graph and produces structured findings — no LLM, no side effects.
| Analyzer | Category | What It Finds |
|---|---|---|
| large_files | size | Files over configurable line thresholds |
| circular_deps | architecture | Import cycles (Tarjan's SCC algorithm) |
| misplaced_imports | architecture | Cross-module dependency bottlenecks |
| hub_bottlenecks | architecture | Files with disproportionate fan-in (z-score outliers) |
| dead_code | quality | Files with zero importers that aren't entry points |
| duplicate_logic | quality | Files with suspiciously similar summaries (Jaccard) |
| tech_debt | quality | Aggregated tech debt from epistemic enrichment |
| staleness | quality | Enrichments that are out of date |
| test_coverage | testing | Source files with no associated test file |
| naming_consistency | naming | Language-specific naming convention violations |
| api_surface | coverage | Public symbols missing docstrings |
Generated Reports
When you run an audit with synthesize: true (via the MCPprep_audit tool or the REST API), an LLM generates 5 markdown documents from the findings:
- AUDIT_SUMMARY — Health grade (A–F), critical findings, top recommendations
- ARCHITECTURE_ANALYSIS — Module dependency flow, bottlenecks, boundary violations
- GAP_ANALYSIS — Misplaced concerns, duplicated logic, missing abstractions
- COMPONENT_INVENTORY — Every file with purpose, module, summary, in-degree
- TECH_DEBT_REPORT — Debt items by module with remediation roadmap
Output Location
All audit output is stored inside the project's index directory:
# Standalone mode (default):
~/.local/share/sourceprep/projects/{project-id}/audit/
├── findings.json # Raw structured findings
├── audit_manifest.json # Run metadata (timestamps, counts)
├── AUDIT_SUMMARY.md # LLM-generated (if synthesized)
├── ARCHITECTURE_ANALYSIS.md
├── GAP_ANALYSIS.md
├── COMPONENT_INVENTORY.md
└── TECH_DEBT_REPORT.md
# Embedded mode:
/path/to/project/.sourceprep/audit/
└── (same files)These files are served via the audit REST API (GET /projects/{id}/audit/reports, GET /projects/{id}/audit/report/{name}) and through the MCPprep_audit tool with action report, so your AI tools can retrieve them by name. (In embedded mode the .sourceprep/audit/ directory is excluded from the walker, so reports are not surfaced via prep_search.)
Settings
Audit behavior is configurable via the audit_config section inui_config.json (edit the file directly; thresholds are read at audit-run time):
| Setting | Default | Description |
|---|---|---|
| large_file_threshold_bytes | 80,000 | File size for "critical" severity (~2000 lines) |
| large_file_warning_bytes | 40,000 | File size for "warning" severity (~1000 lines) |
| hub_z_threshold | 2.0 | Z-score for hub bottleneck detection |
| similarity_threshold | 0.65 | Jaccard threshold for duplicate logic detection |
prep opportunities --format sarif | json | csv | ai_promptMCP: prep_auditLive preview: Opportunities panel consolidating all improvement items with filters and Pi Agent status.
Pipeline Connection
AutoAudit is a finalize pipeline stage. The connection:
- The enrichment pipeline runs to completion and produces trace_nodes, trace_augmented, trace_epistemic, trace_modules, and atlas.json.
- You trigger an audit (via the MCP
prep_audittool, the REST stage-run endpoint, or the dashboard Audit panel) when you want insights. The audit reads all that data and produces findings + reports. - If the per-project
auto_config.finalizeis set toauto, the audit stage runs automatically (Tier 1 + Tier 2) when deep enrichment completes, as part of the finalize group. - Audit reports are served via the audit REST API (GET
/projects/{id}/audit/reports, GET/projects/{id}/audit/report/{name}) and through the MCPprep_audittool with actionreport, so your AI tools can retrieve them by name.
