Writings

Practical notes on real AI systems.

What works, what breaks, and what it costs — published across two newsletters, both written by Nishant Sinha.

Also on Engineering Agents OffNote Labs Newsletter YouTube

Featured
All writings
2026-10-04
Search · Tutorial · 11 steps
Can Agent Build Search for you?
One working session: a coding agent builds product search over the WANDS dataset, and plain grep ranking holds its own against BM25 and vector search.
2026-10-01
Harness
Learn to manage your own context
Context Language Models let the model edit its own context as a file, instead of the harness summarizing on fixed rules: 11.4% more accurate with 21.5% less compute.
2026-09-22
Voice
OpenAI GPT-Live-1 vs Nemotron VoiceChat
Two full duplex voice models, two ways to split the job: keep the conversation live while slow work runs elsewhere, or keep tool calls inside one model.
2026-09-17
The eval gap: why your RAG demo works and your RAG system doesn't
A demo only has to work on the five questions you tried. Production has to work on the five thousand you didn't. Here's the evaluation loop that closes that gap.
2026-09-15
Harness
Agent-driven performance optimization of software systems
Performance work is detective work, the same loop as agentic search. Agents now do it well; the hard part is getting their changes accepted by people.
2026-08-07
Harness
Emergent design: embracing iteration in building software
Coding agents surface corner cases before they become bugs. The scale of building has exploded, but the iterative flavor has not changed.
2026-07-28
Search
Scaling agentic search by merging many parallel queries
Databricks fires many query rewrites in parallel for recall, then merges hundreds of ranked lists around pivot documents for precision.
2026-05-05
Search
Why RAG teams waste months trying the wrong fixes
Chunking, reranking, hybrid search, HyDE: tried one after another, they are expensive guessing. First find out which failure you are fixing.
2026-04-22
Search
Vector vs vectorless RAG is the wrong debate
Most RAG debate compares implementations, not retrieval strategies. Both vector and vectorless search try to do the same thing: localize the query to a small part of the corpus.
2026-04-18
Harness
The Claude Code leak has an unexpected side-effect
A 46-page paper reconstructs Claude Code's design. The core is a simple loop; the real engineering is the harness around it.
2023-05-09
Search
Why your GPT + vector search RAG demo won't make it to production
Embed, store, fetch the nearest vectors, hand them to the LLM. It makes a convincing demo, but it skips decades of what information retrieval learned the hard way.
Have a system that needs a second opinion?