Topics
Browse by subject
Picking a model
Which system to actually use, and whether there is a fair way to compare them.
Vendor claims
What to trust when a lab reports on the quality of its own model.
Prompting
Whether prompt wording changes results, and which techniques survive scrutiny.
Agents
Running models in loops: tool use, self-correction, and where the loop stops helping.
Context engineering
How what you put in the context window changes what the model gives back.
Benchmarks and evaluation
What leaderboards and benchmark scores do, and do not, actually measure.
Reasoning and output levers
The knobs that change reasoning depth and answer accuracy at inference time.
RAG vs fine-tuning
When to retrieve, when to fine-tune, and when to just write a better prompt.
Deployment and product surface
Why the same underlying model feels different inside every app that wraps it.
Teams and roles
How AI changes who does the work, starting with the Forward Deployed Engineer.
Spec-driven development
Writing the specification first, and where the evidence for it actually is.
Model behavior
Personality, consistency, and how models actually behave across prompts and runs.
Hallucination
When and why models make things up, how often it happens, and what actually reduces it.