Entradas de] josecho1969

TIL: Unified Memory Bandwidth is the Real Bottleneck for Local LLMs

When running local AI workloads and large language models (LLMs), raw compute (FLOPS) is often secondary to unified memory bandwidth and capacity. Without sufficient VRAM or unified memory, running 70B+ parameter models at full precision becomes practically impossible without severe quantization, complex multi-GPU setups, or cloud dependencies. I recently read an excellent analysis by Kunal […]

TIL: Knowledge Graph Retrieval – Stop Codebase Re-Reads with Graphify & CodeGraph

Every time an AI coding agent starts a new session, it performs costly file searches and full-file reads («grep-and-read archaeology») to understand your codebase architecture. This process wastes thousands of context tokens on structure discovery alone. The Solution: Local Codebase Indexing & Retrieval Instead of forcing the model to re-read files from scratch, pre-index your […]

TIL: RTK (Rust Token Killer) – 80% Token Reduction for Agent CLI Logs

When AI agents run terminal commands (npm test, git log, docker build, or pytest), the console outputs hundreds of lines of noise, progress bars, and redundant stack traces directly into the model’s context window. This noise quickly consumes thousands of input tokens and triggers context rot. What is RTK? RTK (Rust Token Killer) is an […]

Context Engineering

Every developer working with AI coding agents encounters the same issue: as conversation history grows, reasoning quality degrades and answers become unreliable. This decline is known as Context Rot. Context engineering isn’t about fitting more into the context window—it’s about ruthlessly controlling what stays in it. What is Context Engineering? Context Engineering is the practice […]