Articles, insights, and practical guides on integrating artificial intelligence into software development. Focused on LLMs, terminal-native agents like Claude Code, and building modern AI-assisted engineering workflows.

TIL: Unified Memory Bandwidth is the Real Bottleneck for Local LLMs

When running local AI workloads and large language models (LLMs), raw compute (FLOPS) is often secondary to unified memory bandwidth and capacity.

Without sufficient VRAM or unified memory, running 70B+ parameter models at full precision becomes practically impossible without severe quantization, complex multi-GPU setups, or cloud dependencies.

I recently read an excellent analysis by Kunal Ganglani on how Apple’s M5 Max chip architecture impacts local AI execution:

🔗 Reading: Apple’s M5 Max Just Made the Case for Local AI Development

Key Takeaways for AI Engineers:

  • The Memory Wall: Token generation in local LLM inference is strictly memory-bandwidth bound.
  • Unified Memory vs. VRAM: Having up to 128GB of unified memory at 614 GB/s enables loading massive models locally without splitting across multiple GPUs or taking severe quality hits via quantization.
  • Frictionless Prototyping: Compress the gap between testing a model locally and deploying agentic workflows without cloud costs or DevOps overhead.

A must-read breakdown if you are designing local fallback environments for AI agents.

Context Engineering

Every developer working with AI coding agents encounters the same issue: as conversation history grows, reasoning quality degrades and answers become unreliable. This decline is known as Context Rot.

Context engineering isn’t about fitting more into the context window—it’s about ruthlessly controlling what stays in it.


What is Context Engineering?

Context Engineering is the practice of actively managing, filtering, and structuring the information provided to Large Language Models (LLMs) during interactive agentic sessions.

Rather than treating the context window as an endless dumping ground for terminal logs, full codebase files, and long conversation turns, context engineering applies strict operational filters to ensure the model receives only actionable, high-density signal.


The 10-Tactic Context Engineering Framework

Layer 1: Context Window & Session Hygiene

01. Keep the Context Window Short (The Absolute Token Rule)
Focus on absolute token counts, not percentages: A 1M token context window does not mean quality remains intact at 400k tokens.

  • The Sweet Spot (<100k tokens): Keep working sessions under 100k tokens for peak reasoning precision.
  • Warning Zone (100k–200k tokens): Clear and start fresh soon. Beyond 200k tokens is an unpredictable bet.
  • Key Metric: Benchmarks show that 18 out of 18 leading models suffer measurable reasoning degradation well before reaching their nominal context limits.

02. Clear Early, Clear Often (Strict Session Hygiene)
Enforce a strict «One Task, One Session» workflow. Never keep «kitchen sink» sessions running indefinitely.

  • The Rule: If you have corrected the agent twice on the same mistake or if the token counter crosses 100k, wipe the session and re-prompt with a tighter context.
  • Clearing vs. Compaction: Instead of relying on automatic compaction (which acts as a late-stage airbag), manually write current progress to a handoff.md file and resume from a fresh session.

03. Short & Sharp Prompts (Precision Input Curation)
Point the agent to explicit file references rather than pasting raw code blocks into the chat prompt.

  • Scope Every Read: Instruct the agent to read specific modules (e.g., src/auth/session.ts) rather than entire directories.
  • Front-load Constraints: Place acceptance criteria and hard rules at the top of the prompt where model attention is strongest.
  • Key Metric: Model completion accuracy drops by up to 39% when user intent and instruction constraints arrive in fragmented, multi-part messages.

Layer 2: Environment & Tool Configuration

04. AGENTS.md Diet (Lean Persistent Rules)
Keep system instructions lean and ruthless. Audit your root and local AGENTS.md files constantly.

  • The Audit Rule: Only include instructions that, if removed, would directly cause the agent to make mistakes. Evict project origin stories and generic best practices.
  • Directory-Level Modularity: Split instructions into nested, per-directory AGENTS.md files that load only when the agent works in that scope.
  • Key Metric: Top agentic runtimes (like OpenAI Codex) cap system instruction files at 32 KiB.

05. MCP Tool Loadout Diet
Avoid loading monolithic Model Context Protocol (MCP) server configurations across all environments.

  • On-Demand Schemas: Every connected MCP server injects schema definitions on every turn (~1k tokens per tool). Prefer native CLI tools (gh, psql, curl) which cost zero schema tokens.
  • Project-Scoped Loadouts: Configure MCP tools per project and keep active tools under 30 to avoid confusing model reasoning.

06. Skill Refinement (Lazy-Loaded Agent Skills)
Build highly specialized, modular skills rather than huge catch-all instruction manuals.

  • Progressive Disclosure: Skills load minimal metadata at startup and fetch their full body only when explicitly triggered.
  • Key Metric: Limit skill definitions to <500 lines per SKILL.md. Rely on local scripts for deterministic execution.

Layer 3: Architectural Workflow

07. Plan Files Strategy (External Memory & Handoffs)
Separate architectural research, execution planning, and code implementation into discrete phases. State stored in plan.md, specs.md, or todo.md survives session wipes.

  • The Rule: Use a 3-step handoff pipeline: research.md → plan.md → implementation. Start a fresh, clean chat session for each stage.
  • Key Metric: A clean 200-line implementation plan effectively replaces up to 150,000 tokens of noisy search and trial-and-error chat history.

Layer 4: The CLI & Retrieval Optimization Stack

08. Terse Output (Caveman)
Enforce concise, low-prose response formatting to minimize output token generation costs and accelerate turnarounds.

  • The Tool: Apply /caveman full configurations or run compression passes on local memory files.
  • Key Metric: Delivers a 25% to 50% net token reduction on output generation.

09. Mute the Shell (RTK)
Interception layer that strips noise, progress bars, and redundant stack traces from CLI tool execution.

  • The Tool: Install via brew install rtk and initialize with rtk init -g ..
  • Key Metric: Cuts raw terminal output noise by up to 80% per session.

💡 Enterprise Adoption Note:

Terminal output ≠ Total bill: RTK cuts shell noise by ~80%, but system prompts, file reads, and history still make up the bulk of your context.

Debugging risk: Aggressive filtering can drop subtle error lines. Always keep tee mode enabled for full local logging.

Supply Chain & Security: A third-party binary reading command logs requires client approval and security review before corporate adoption.

💡 For an enterprise adoption breakdown and security caveats on RTK, check out my recent TIL: Enterprise Trade-offs of CLI Filters like RTK.

10. Kill Re-Reads (Graphify & CodeGraph)
Eliminate costly «grep-and-read» codebase discovery at the start of every agent session.

  • The Tools: Pre-index your repository structure into queryable local knowledge graphs using Graphify or CodeGraph.
  • Key Metric: Saves between 18% and 90% of input tokens by converting multi-file scans into targeted graph queries.

Quick Reference Summary

  • 01. Window Size: Keep active context <100k tokens (degradation hits well before max limit).
  • 02. Session Hygiene: Reset session after 2 failed fixes or 100k tokens; write progress to handoff.md.
  • 03. Prompting: Point to files, don’t paste code (prevents 39% accuracy drop).
  • 04. AGENTS.md: Remove unnecessary rules per-line; use per-directory files (stay under 32 KiB cap).
  • 05. MCP Loadout: Use per-project tools and CLI native tools (>30 active tools confuses models).
  • 06. Skills: Limit skill files to <500 lines per SKILL.md with deterministic script execution.
  • 07. Workflows: Use research.md → plan.md → code (200-line plan replaces 150k tokens).
  • 08. Caveman: Use terse mode for 25%–50% output token savings.
  • 09. RTK: Filter shell output with RTK proxy for 80% terminal noise reduction.
  • 10. Graphify / CodeGraph: Index codebase structure to save 18%–90% on discovery re-reads.

When Claude Code Lost My Commit: How VS Code Local History Saved the Day

AI agents in the terminal are powerful, but they aren’t bulletproof. Here is how I recovered lost changes using VS Code’s built-in tools.

I have been using Claude Code directly in my terminal for several months as part of my daily Full Stack development workflow. It has significantly boosted my speed when refactoring code, writing unit tests, and building feature boilerplates.

However, working with autonomous AI agents comes with a learning curve. A few days ago, while prompting Claude Code to execute a refactoring task, something went wrong during a git operation, and a recent commit with unpushed changes seemed to vanish.

If you use AI assistants in your terminal, you will eventually face a moment where the agent overrules or overwrites something unexpectedly. Here is what happened and how I recovered my code in under two minutes.

The Incident: A Missing Commit

While working on a feature, I asked Claude Code to perform a multi-file refactor and handle the commit process. Due to a context mismatch in the terminal execution, the recent local commit was detached, leaving my working directory in an unexpected state without the updated changes.

My immediate reaction was to check git reflog, but since the changes hadn’t been fully tracked in the branch head yet, the standard Git safety net wasn’t enough.

That’s when Claude Code itself suggested a feature built right into Visual Studio Code: Local History.

The Solution: VS Code Local History

VS Code automatically tracks local revisions of your files every time you save them, independent of Git commits or stashes. This timeline acts as an extra safety net when local files are overwritten by external scripts, CLI commands, or AI agents.

Here is how to access and restore your files:

  1. Open the file that was overwritten or lost in VS Code.
  2. Open the Explorer panel on the left sidebar.
  3. Scroll down to the Timeline section.
  4. Click through the chronological list of file saves to compare differences (diff view).
  5. Right-click the snapshot prior to the incident and select Restore Contents.
VS Code Explorer panel showing the Timeline section used to recover lost file history after a Claude Code terminal incident.
Restoring a previous file version using the Timeline panel in Visual Studio Code.

Tip: You can also right-click any timestamp in the Timeline panel to compare that specific version directly against your current file state.

Key Takeaways for AI-Assisted Workflows

Working with AI tools in the CLI requires adopting a few defensive habits:

  1. Commit frequently: Make small, granular Git commits before letting an AI agent execute complex multi-file refactors.
  2. Don’t rely solely on Git: Keep tools like VS Code’s Timeline / Local History in mind for local file recovery.
  3. Review agent actions: Always inspect the diffs before running destructive git commands or script executions suggested by terminal agents.

Have you ever lost code while working with an AI CLI tool? How do you manage safety nets in your developer workflow?