TIL / Recursos de Aprendizaje

TIL: Unified Memory Bandwidth is the Real Bottleneck for Local LLMs

When running local AI workloads and large language models (LLMs), raw compute (FLOPS) is often secondary to unified memory bandwidth and capacity.

Without sufficient VRAM or unified memory, running 70B+ parameter models at full precision becomes practically impossible without severe quantization, complex multi-GPU setups, or cloud dependencies.

I recently read an excellent analysis by Kunal Ganglani on how Apple’s M5 Max chip architecture impacts local AI execution:

🔗 Reading: Apple’s M5 Max Just Made the Case for Local AI Development

Key Takeaways for AI Engineers:

  • The Memory Wall: Token generation in local LLM inference is strictly memory-bandwidth bound.
  • Unified Memory vs. VRAM: Having up to 128GB of unified memory at 614 GB/s enables loading massive models locally without splitting across multiple GPUs or taking severe quality hits via quantization.
  • Frictionless Prototyping: Compress the gap between testing a model locally and deploying agentic workflows without cloud costs or DevOps overhead.

A must-read breakdown if you are designing local fallback environments for AI agents.

TIL: Knowledge Graph Retrieval – Stop Codebase Re-Reads with Graphify & CodeGraph

Every time an AI coding agent starts a new session, it performs costly file searches and full-file reads («grep-and-read archaeology») to understand your codebase architecture. This process wastes thousands of context tokens on structure discovery alone.

The Solution: Local Codebase Indexing & Retrieval

Instead of forcing the model to re-read files from scratch, pre-index your repository locally using knowledge graph tools. This allows the agent to query structural relationships (functions, classes, schemas, imports) directly via scoped queries.


Two High-Impact Open-Source Tools

1. Graphify (by Safi Shamsi)

Parses source code into a queryable knowledge graph using tree-sitter (supporting 30+ languages). Check out the Graphify repository on GitHub.

  • Key Feature: Integrates documentation and PDFs via local LLM extraction. Connects natively with 24+ harnesses via MCP for team sharing.
  • Impact: Up to 71.5x fewer tokens consumed per codebase architectural query.
  • Quick Start: uv tool install graphifyy && graphify install

2. CodeGraph (by Colby McHenry)

A zero-maintenance, background code indexing tool stored in a local SQLite database with full-text search. Explore the CodeGraph repository on GitHub or find the npm package.

  • Key Feature: Automatically watches file changes and re-syncs continuously in the background—no API keys required.
  • Impact: Delivers 23–64% token savings and 58% fewer tool calls during repo navigation.
  • Quick Start: npm i -g @colbymchenry/codegraph

Rule of Thumb: Use Graphify if you need multimodal context, deep architectural relationships, or team MCP sharing. Use CodeGraph for set-it-and-forget-it local navigation that stays auto-synced.

TIL: RTK (Rust Token Killer) – 80% Token Reduction for Agent CLI Logs

When AI agents run terminal commands (npm test, git log, docker build, or pytest), the console outputs hundreds of lines of noise, progress bars, and redundant stack traces directly into the model’s context window.

This noise quickly consumes thousands of input tokens and triggers context rot.

What is RTK?

RTK (Rust Token Killer) is an open-source proxy filter (Apache 2.0) that sits between your terminal and the AI agent to intercept and compress command output before it enters the context window.

Key Technical Takeaways & Impact:

  • Terminal Interception: Automatically filters output for over 100 common CLI commands (git, pnpm, docker, cargo, pytest).
  • Noise Reduction: Deduplicates repeated error traces, strips out loading/progress bars, and keeps only critical failure points.
  • Tee Mode (Zero Loss): If a command fails, full uncompressed log files are saved to disk for human inspection, while sending only the essential error summary to the AI.
  • Measured Token Savings: Real-world benchmark tests show session log reductions from 111,000 tokens down to 23,000 tokens (an ~80% reduction in context overhead).
  • Easy Setup: Installable via Homebrew (brew install rtk) and integrates natively into tools like Claude Code, Cursor CLI, and GitHub Copilot.

Bottom Line: If you only install one CLI optimization tool for AI development, make it RTK. It is the fastest way to slash token usage and prevent agent context rot during terminal execution.

This analysis is part of the broader framework covered in Context Engineering Tactics: A 10-Tactic Framework for AI Agents.