>_ JLVBCoop FULL STACK & AI ENGINEERING >_ JLVBCoop FULL STACK & AI ENGINEERING
  • AI Guide
  • TIL
  • Blog
  • About
  • Click to open the search input field Click to open the search input field Buscar
  • Menú Menú

Technical Blog & AI Engineering

Real-world insights on Full Stack development, AI-assisted workflows, and software architecture.

Note: I’m a native Spanish-speaking Full Stack Developer writing in English to enhance my technical communication skills. Feedback on code, architecture, or grammar is always welcome!

TIL: Unified Memory Bandwidth is the Real Bottleneck for Local LLMs

4 de octubre de 2026/en TIL, AI Engineering, Hardware/por josecho1969

When running local AI workloads and large language models (LLMs), raw compute (FLOPS) is often secondary to unified memory bandwidth and capacity.

Without sufficient VRAM or unified memory, running 70B+ parameter models at full precision becomes practically impossible without severe quantization, complex multi-GPU setups, or cloud dependencies.

I recently read an excellent analysis by Kunal Ganglani on how Apple’s M5 Max chip architecture impacts local AI execution:

🔗 Reading: Apple’s M5 Max Just Made the Case for Local AI Development

Key Takeaways for AI Engineers:

  • The Memory Wall: Token generation in local LLM inference is strictly memory-bandwidth bound.
  • Unified Memory vs. VRAM: Having up to 128GB of unified memory at 614 GB/s enables loading massive models locally without splitting across multiple GPUs or taking severe quality hits via quantization.
  • Frictionless Prototyping: Compress the gap between testing a model locally and deploying agentic workflows without cloud costs or DevOps overhead.

A must-read breakdown if you are designing local fallback environments for AI agents.

https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg 0 0 josecho1969 https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg josecho19692026-10-04 09:04:242026-10-04 09:05:55TIL: Unified Memory Bandwidth is the Real Bottleneck for Local LLMs

TIL: Knowledge Graph Retrieval – Stop Codebase Re-Reads with Graphify & CodeGraph

3 de octubre de 2026/en TIL/por josecho1969

Every time an AI coding agent starts a new session, it performs costly file searches and full-file reads («grep-and-read archaeology») to understand your codebase architecture. This process wastes thousands of context tokens on structure discovery alone.

The Solution: Local Codebase Indexing & Retrieval

Instead of forcing the model to re-read files from scratch, pre-index your repository locally using knowledge graph tools. This allows the agent to query structural relationships (functions, classes, schemas, imports) directly via scoped queries.


Two High-Impact Open-Source Tools

1. Graphify (by Safi Shamsi)

Parses source code into a queryable knowledge graph using tree-sitter (supporting 30+ languages). Check out the Graphify repository on GitHub.

  • Key Feature: Integrates documentation and PDFs via local LLM extraction. Connects natively with 24+ harnesses via MCP for team sharing.
  • Impact: Up to 71.5x fewer tokens consumed per codebase architectural query.
  • Quick Start: uv tool install graphifyy && graphify install

2. CodeGraph (by Colby McHenry)

A zero-maintenance, background code indexing tool stored in a local SQLite database with full-text search. Explore the CodeGraph repository on GitHub or find the npm package.

  • Key Feature: Automatically watches file changes and re-syncs continuously in the background—no API keys required.
  • Impact: Delivers 23–64% token savings and 58% fewer tool calls during repo navigation.
  • Quick Start: npm i -g @colbymchenry/codegraph

Rule of Thumb: Use Graphify if you need multimodal context, deep architectural relationships, or team MCP sharing. Use CodeGraph for set-it-and-forget-it local navigation that stays auto-synced.

https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg 0 0 josecho1969 https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg josecho19692026-10-03 07:46:312026-10-03 07:56:29TIL: Knowledge Graph Retrieval – Stop Codebase Re-Reads with Graphify & CodeGraph

TIL: RTK (Rust Token Killer) – 80% Token Reduction for Agent CLI Logs

3 de octubre de 2026/en TIL/por josecho1969

When AI agents run terminal commands (npm test, git log, docker build, or pytest), the console outputs hundreds of lines of noise, progress bars, and redundant stack traces directly into the model’s context window.

This noise quickly consumes thousands of input tokens and triggers context rot.

What is RTK?

RTK (Rust Token Killer) is an open-source proxy filter (Apache 2.0) that sits between your terminal and the AI agent to intercept and compress command output before it enters the context window.

Key Technical Takeaways & Impact:

  • Terminal Interception: Automatically filters output for over 100 common CLI commands (git, pnpm, docker, cargo, pytest).
  • Noise Reduction: Deduplicates repeated error traces, strips out loading/progress bars, and keeps only critical failure points.
  • Tee Mode (Zero Loss): If a command fails, full uncompressed log files are saved to disk for human inspection, while sending only the essential error summary to the AI.
  • Measured Token Savings: Real-world benchmark tests show session log reductions from 111,000 tokens down to 23,000 tokens (an ~80% reduction in context overhead).
  • Easy Setup: Installable via Homebrew (brew install rtk) and integrates natively into tools like Claude Code, Cursor CLI, and GitHub Copilot.

Bottom Line: If you only install one CLI optimization tool for AI development, make it RTK. It is the fastest way to slash token usage and prevent agent context rot during terminal execution.

This analysis is part of the broader framework covered in Context Engineering Tactics: A 10-Tactic Framework for AI Agents.

https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg 0 0 josecho1969 https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg josecho19692026-10-03 07:32:042026-10-05 15:41:04TIL: RTK (Rust Token Killer) – 80% Token Reduction for Agent CLI Logs

Context Engineering

3 de octubre de 2026/en AI Engineering/por josecho1969

Every developer working with AI coding agents encounters the same issue: as conversation history grows, reasoning quality degrades and answers become unreliable. This decline is known as Context Rot.

Context engineering isn’t about fitting more into the context window—it’s about ruthlessly controlling what stays in it.


What is Context Engineering?

Context Engineering is the practice of actively managing, filtering, and structuring the information provided to Large Language Models (LLMs) during interactive agentic sessions.

Rather than treating the context window as an endless dumping ground for terminal logs, full codebase files, and long conversation turns, context engineering applies strict operational filters to ensure the model receives only actionable, high-density signal.


The 10-Tactic Context Engineering Framework

Layer 1: Context Window & Session Hygiene

01. Keep the Context Window Short (The Absolute Token Rule)
Focus on absolute token counts, not percentages: A 1M token context window does not mean quality remains intact at 400k tokens.

  • The Sweet Spot (<100k tokens): Keep working sessions under 100k tokens for peak reasoning precision.
  • Warning Zone (100k–200k tokens): Clear and start fresh soon. Beyond 200k tokens is an unpredictable bet.
  • Key Metric: Benchmarks show that 18 out of 18 leading models suffer measurable reasoning degradation well before reaching their nominal context limits.

02. Clear Early, Clear Often (Strict Session Hygiene)
Enforce a strict «One Task, One Session» workflow. Never keep «kitchen sink» sessions running indefinitely.

  • The Rule: If you have corrected the agent twice on the same mistake or if the token counter crosses 100k, wipe the session and re-prompt with a tighter context.
  • Clearing vs. Compaction: Instead of relying on automatic compaction (which acts as a late-stage airbag), manually write current progress to a handoff.md file and resume from a fresh session.

03. Short & Sharp Prompts (Precision Input Curation)
Point the agent to explicit file references rather than pasting raw code blocks into the chat prompt.

  • Scope Every Read: Instruct the agent to read specific modules (e.g., src/auth/session.ts) rather than entire directories.
  • Front-load Constraints: Place acceptance criteria and hard rules at the top of the prompt where model attention is strongest.
  • Key Metric: Model completion accuracy drops by up to 39% when user intent and instruction constraints arrive in fragmented, multi-part messages.

Layer 2: Environment & Tool Configuration

04. AGENTS.md Diet (Lean Persistent Rules)
Keep system instructions lean and ruthless. Audit your root and local AGENTS.md files constantly.

  • The Audit Rule: Only include instructions that, if removed, would directly cause the agent to make mistakes. Evict project origin stories and generic best practices.
  • Directory-Level Modularity: Split instructions into nested, per-directory AGENTS.md files that load only when the agent works in that scope.
  • Key Metric: Top agentic runtimes (like OpenAI Codex) cap system instruction files at 32 KiB.

05. MCP Tool Loadout Diet
Avoid loading monolithic Model Context Protocol (MCP) server configurations across all environments.

  • On-Demand Schemas: Every connected MCP server injects schema definitions on every turn (~1k tokens per tool). Prefer native CLI tools (gh, psql, curl) which cost zero schema tokens.
  • Project-Scoped Loadouts: Configure MCP tools per project and keep active tools under 30 to avoid confusing model reasoning.

06. Skill Refinement (Lazy-Loaded Agent Skills)
Build highly specialized, modular skills rather than huge catch-all instruction manuals.

  • Progressive Disclosure: Skills load minimal metadata at startup and fetch their full body only when explicitly triggered.
  • Key Metric: Limit skill definitions to <500 lines per SKILL.md. Rely on local scripts for deterministic execution.

Layer 3: Architectural Workflow

07. Plan Files Strategy (External Memory & Handoffs)
Separate architectural research, execution planning, and code implementation into discrete phases. State stored in plan.md, specs.md, or todo.md survives session wipes.

  • The Rule: Use a 3-step handoff pipeline: research.md → plan.md → implementation. Start a fresh, clean chat session for each stage.
  • Key Metric: A clean 200-line implementation plan effectively replaces up to 150,000 tokens of noisy search and trial-and-error chat history.

Layer 4: The CLI & Retrieval Optimization Stack

08. Terse Output (Caveman)
Enforce concise, low-prose response formatting to minimize output token generation costs and accelerate turnarounds.

  • The Tool: Apply /caveman full configurations or run compression passes on local memory files.
  • Key Metric: Delivers a 25% to 50% net token reduction on output generation.

09. Mute the Shell (RTK)
Interception layer that strips noise, progress bars, and redundant stack traces from CLI tool execution.

  • The Tool: Install via brew install rtk and initialize with rtk init -g ..
  • Key Metric: Cuts raw terminal output noise by up to 80% per session.

💡 Enterprise Adoption Note:

Terminal output ≠ Total bill: RTK cuts shell noise by ~80%, but system prompts, file reads, and history still make up the bulk of your context.

Debugging risk: Aggressive filtering can drop subtle error lines. Always keep tee mode enabled for full local logging.

Supply Chain & Security: A third-party binary reading command logs requires client approval and security review before corporate adoption.

💡 For an enterprise adoption breakdown and security caveats on RTK, check out my recent TIL: Enterprise Trade-offs of CLI Filters like RTK.

10. Kill Re-Reads (Graphify & CodeGraph)
Eliminate costly «grep-and-read» codebase discovery at the start of every agent session.

  • The Tools: Pre-index your repository structure into queryable local knowledge graphs using Graphify or CodeGraph.
  • Key Metric: Saves between 18% and 90% of input tokens by converting multi-file scans into targeted graph queries.

Quick Reference Summary

  • 01. Window Size: Keep active context <100k tokens (degradation hits well before max limit).
  • 02. Session Hygiene: Reset session after 2 failed fixes or 100k tokens; write progress to handoff.md.
  • 03. Prompting: Point to files, don’t paste code (prevents 39% accuracy drop).
  • 04. AGENTS.md: Remove unnecessary rules per-line; use per-directory files (stay under 32 KiB cap).
  • 05. MCP Loadout: Use per-project tools and CLI native tools (>30 active tools confuses models).
  • 06. Skills: Limit skill files to <500 lines per SKILL.md with deterministic script execution.
  • 07. Workflows: Use research.md → plan.md → code (200-line plan replaces 150k tokens).
  • 08. Caveman: Use terse mode for 25%–50% output token savings.
  • 09. RTK: Filter shell output with RTK proxy for 80% terminal noise reduction.
  • 10. Graphify / CodeGraph: Index codebase structure to save 18%–90% on discovery re-reads.
https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg 0 0 josecho1969 https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg josecho19692026-10-03 07:13:102026-10-05 15:39:59Context Engineering

TIL: Caveman – Lightweight Context & File Packing for LLMs

3 de octubre de 2026/en TIL/por josecho1969
Leer más
https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg 0 0 josecho1969 https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg josecho19692026-10-03 06:59:422026-10-03 07:15:27TIL: Caveman – Lightweight Context & File Packing for LLMs

My First Steps with AI-SDLC: Spec-Driven Development (with AI)

1 de octubre de 2026/en AI Engineering/por josecho1969
Leer más
https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg 0 0 josecho1969 https://jlvbcoop.com/wp-content/uploads/2026/09/jlvbcoopLogo.svg josecho19692026-10-01 21:46:502026-10-01 22:01:55My First Steps with AI-SDLC: Spec-Driven Development (with AI)
Página 1 de 212

Search Search

Entradas recientes

  • TIL: Unified Memory Bandwidth is the Real Bottleneck for Local LLMs
  • TIL: Knowledge Graph Retrieval – Stop Codebase Re-Reads with Graphify & CodeGraph
  • TIL: RTK (Rust Token Killer) – 80% Token Reduction for Agent CLI Logs
  • Context Engineering
  • TIL: Caveman – Lightweight Context & File Packing for LLMs

Archivos

  • octubre 2026
  • septiembre 2026

Categorías

  • AI Engineering
  • Hardware
  • TIL

JLVBCoop

Full Stack & AI Engineer with 15+ years of experience. Building robust software with Java, React, and modern AI workflows.

Categorías

  • AI Engineering
  • Hardware
  • TIL

Connect

jlvbalsa@gmail.com

Linkedin

Entradas recientes

  • TIL: Unified Memory Bandwidth is the Real Bottleneck for Local LLMs
  • TIL: Knowledge Graph Retrieval – Stop Codebase Re-Reads with Graphify & CodeGraph
  • TIL: RTK (Rust Token Killer) – 80% Token Reduction for Agent CLI Logs
  • Context Engineering
  • TIL: Caveman – Lightweight Context & File Packing for LLMs
© Copyright - jlvbcoop. All rights reserved. - powered by Enfold WordPress Theme
Desplazarse hacia arriba Desplazarse hacia arriba Desplazarse hacia arriba