Ollama that sizes itself to your GPU
Native Ollama API, VRAM detection on NVIDIA and Apple Silicon, and a num_ctx that shrinks on memory pressure instead of hanging.
codetrace set-ctx --backoff 0.75CodeTrace AI turns your repository into a deterministic call graph and makes the agent prove every claim — verified blast radius and exact file:line citations before any edit. Parsing, embeddings and the graph run on your machine; bring any LLM, or pair it with Ollama and nothing leaves it at all.

Real session · shown at 2× speed · qwen3.5:4b via Ollama on a 6 GB laptop GPU · nothing left the machine
The biggest release since launch: roughly 4,100 lines of new code across the agent loop, token manager, MCP server and indexer.
Native Ollama API, VRAM detection on NVIDIA and Apple Silicon, and a num_ctx that shrinks on memory pressure instead of hanging.
codetrace set-ctx --backoff 0.75Stdio server your IDE launches, with one entry per project in .mcp.json, .cursor/mcp.json and .vscode/mcp.json.
codetrace register-mcp .The custom provider speaks both the OpenAI and Anthropic protocols — DeepSeek, vLLM, LM Studio, corporate gateways. Keys optional for local servers.
codetrace config # → customEvery structural claim is graded CONFIRMED, INFERRED or UNRESOLVED and carries a live file:line citation. Output is normalized across providers.
codetrace chatAPI-key config written atomically and owner-only, stricter project-path checks, writes outside the project refused before you're asked.
~/.codetrace/config.json (0600)git_diff no longer returns empty, re-indexing keeps inbound call edges, timeouts show a real error, and CPU-only machines stop running out of memory while re-ranking.
pytest # 55 passingAI coding agents break production systems because they operate on unverified assumptions. CodeTrace AI enforces a formal evidence governance layer at the agent-loop level: every structural assertion must carry live verification.
Established directly by a live AST, symbol relation, or file snapshot tool execution in the active session.
[CONFIRMED] auth/jwt.py:42
def verify(token: str, secret: str) -> dict:
# Established via live get_symbol_relations() call
# Callers: middleware/auth.py:18, api/routes/users.py:88Logical deduction derived from confirmed call patterns and graph topologies — strictly transparent.
[INFERRED] Based on the call pattern between auth/jwt.py and
middleware/auth.py: modifying verify() parameter signature
will cascade an authentication bypass exception across all protected routes.Insufficient or ambiguous evidence — the agent explicitly refuses to guess and specifies what is missing.
[UNRESOLVED] Symbol 'OAuthCallbackHandler' is referenced in routes/auth.py:14
but definition is not found in local index.
Action: Run 'codetrace index --fast' to include external vendor packages.
Applied in the narrowest sequence necessary. Enforced in the agent loop, not just the prompt: no tool evidence, no structural claim.
Hybrid semantic search (BGE + E5 + RRF + FlashRank reranker) retrieves candidate files.
Traverses SQLite + NetworkX call graph to identify callers, callees, and imports.
Reads exact line ranges from indexed DB snapshots with path traversal safeguards.
Calculates transitive blast radius across modules, routes, and test suites.
Compiles evidence-graded response with mandatory file:line citations & output normalization.
file:line citation from tool output. No citation → no claim.Explore CodeTrace AI's real developer terminal interface, hardware detection, dynamic model discovery, and IDE integration live.
CodeTrace AI boots in your terminal or IDE with automatic hardware resource detection (CPU cores, RAM, GPU context window), persistent session management, and governed evidence reasoning.

From first clone to confident shipping — everything you and your AI agent need to inspect, query, and refactor code safely.
Ask complex engineering questions in natural language. The agent executes hybrid search, navigates AST symbols, and reads file ranges to formulate grounded answers with verified file:line citations.
# Query: "Where is UserToken issued and verified?"
auth/jwt.py:42 issues JWT tokens via create_token()
middleware/auth.py:18 verifies tokens via verify()
tests/test_auth.py:64 validates cryptographic signatures
Coverage: 3 references, 6 active callers [CONFIRMED]Maps exact caller/callee relationships across 21 languages using Tree-sitter ASTs. See every downstream function, route, database model, and test suite affected by a planned change.
● auth/jwt.py (Target Symbol: verify)
├── 1-hop: middleware/auth.py:18 (Auth Guard)
├── 1-hop: api/routes/users.py:42 (User Profile)
├── 2-hop: api/routes/admin.py:91 (Admin Dashboard)
└── Test: tests/test_auth.py (14 test cases impacted)Generate an offline, interactive 3D/2D visual graph of your codebase architecture. Features collapsible directory trees, hover symbol inspectors, live filter search, and cross-folder call edges.
$ codetrace visualize
✓ Generated .codetrace/graph_visualization.html
✓ 142 Nodes, 4,891 Edges rendered
✓ Cross-folder dependency linkages mapped
Opening in default browser...Code modifications are proposed as clean unified diffs with built-in path-traversal protection. Nothing is ever written to disk without explicit developer confirmation.
--- a/middleware/auth.py
+++ b/middleware/auth.py
@@ -18,3 +18,4 @@
- token = request.headers.get("Authorization")
+ token = sanitize_bearer(request.headers.get("Authorization"))
+ claims = verify(token, config.SECRET_KEY)
[Apply this diff to disk? (y/n/review)]:Tracks file checksums to only re-parse files that actually changed. Subsequent runs complete in milliseconds even on million-line monolithic repositories.
$ codetrace index .
[Delta Sync] 139 files unchanged (hash match)
[Delta Sync] 3 files modified → re-indexed in 0.38s
[Vector Sync] ChromaDB delta updated successfullyNormalizes headings, list symbols, code-fence aliases (`py` → `python`), and spacing across all models — from local 7B Ollama to frontier cloud LLMs, ensuring a uniform CLI aesthetic.
# output_normalizer.py in agent loop
def normalize_response(raw_text: str) -> str:
# Normalizes code fences, heading levels,
# removes markdown artifacts, enforces uniform CLI styling
return normalized_rich_textFrom raw source files to governed AI reasoning and MCP IDE integration. Every layer is inspectable, modular, and 100% offline-first.
Ingests any local directory or clones a GitHub URL. Computes SHA-256 delta hashes per file to bypass unchanged files during incremental sync.
[Ingestion Engine]
Hashing repository tree...
Checked 142 files via SHA-256
Unmodified: 139 files (cached)
Changed: 3 files scheduled for AST passMulti-layer symbol graph interconnecting Tree-sitter ASTs, BGE/E5 dense embeddings, SQLite relational tables, NetworkX call graphs, and live MCP reasoning sessions.
The AI Architect and external MCP clients invoke these 7 tools autonomously to search, inspect, traverse, and propose changes without guessing.
Executes dense vector embedding search + sparse BM25 retrieval merged via Reciprocal Rank Fusion (RRF) and scored with local FlashRank neural reranking.
| Param | Type | Description |
|---|---|---|
| query | string | Natural language query describing symbol, behavior, or feature |
| limit | int | Maximum precision candidate snippets to return (default: 5) |
[
{
"file": "middleware/rate_limiter.py",
"lines": "24-58",
"score": 0.942,
"snippet": "class TokenBucketLimiter:\n def allow_request(self, key): ..."
},
{
"file": "config/limits.yaml",
"lines": "1-15",
"score": 0.887,
"snippet": "rate_limits:\n api_v1: 100/min\n admin: 500/min"
}
]codetrace init automatically registers the MCP server in Cursor, Claude Code, and VS Code. Your favorite editor instantly gains access to all 7 tools for in-editor AI assistance.
CodeTrace automatically injects the MCP server configuration into your Cursor settings during `codetrace init`.
~/.cursor/mcp.json{
"mcpServers": {
"codetrace": {
"command": "python",
"args": ["-m", "codetrace_mcp.server", "--project", "/path/to/your/project"]
}
}
}
CodeTrace ships with six native providers out of the box plus a universal custom option for self-hosted vLLM, LM Studio, DeepSeek, or corporate proxies.
Test your endpoint settings. CodeTrace prompts for these during codetrace config and stores them at ~/.codetrace/config.json.
{
"provider": "custom",
"api_style": "openai",
"base_url": "https://api.deepseek.com/v1",
"model_name": "deepseek-chat"
}
CodeTrace automatically detects your GPU VRAM, sizes num_ctx safely, and backs off on CPU memory pressure so your system never hangs.
Fast inference on ultrabooks and standard laptops.
Optimal balance of speed and deep structural reasoning.
Near frontier-level reasoning entirely on your workstation.
17 parsed with native Tree-sitter grammars into complete symbol + call graphs; 4 config and data formats parsed into structural symbol hierarchies.
Click any file below to inspect its live call-graph dependents, blast radius impact score, and proposed safe diff preview.
--- a/auth/jwt.py
+++ b/auth/jwt.py
@@ -42,4 +42,5 @@
-def verify(token: str, secret: str) -> dict:
- return jwt.decode(token, secret)
+def verify(token: str, secret: str, algorithms: list = ["RS256"]) -> dict:
+ return jwt.decode(token, secret, algorithms=algorithms)
[Awaiting Human Approval]Everything in CodeTrace AI is accessible through intuitive CLI commands with rich shell output, progress indicators, and flags.
One-command setup: configure LLM provider, download embedding models, index repository, and register MCP in Cursor, Claude Code, and VS Code.
--offlineStrict air-gapped mode (blocks external calls)--fastUse smaller embedding models for lower RAM usage--llm <provider>Pre-select provider: anthropic, openai, gemini, groq, openrouter, ollama, customcd /path/to/my-project
codetrace init
# Or air-gapped mode:
codetrace init --offline
CodeTrace AI is actively developed with rapid improvements in deterministic AST parsing, local agent governance, and IDE integration.
The core codetrace-ai package is completely free, MIT licensed, and runs entirely on your hardware. Join our waitlist for first access to the upcoming hosted team engine.
Full-featured local governed code intelligence for individual developers and air-gapped systems.
High-throughput cloud-accelerated repository indexing, multi-repo workspaces, and collaborative team intelligence.