Topics

Browse by subject

Picking a model

Which system to actually use, and whether there is a fair way to compare them.

Vendor claims

What to trust when a lab reports on the quality of its own model.

Prompting

Whether prompt wording changes results, and which techniques survive scrutiny.

Agents

Running models in loops: tool use, self-correction, and where the loop stops helping.

Context engineering

How what you put in the context window changes what the model gives back.

Benchmarks and evaluation

What leaderboards and benchmark scores do, and do not, actually measure.

Reasoning and output levers

The knobs that change reasoning depth and answer accuracy at inference time.

RAG vs fine-tuning

When to retrieve, when to fine-tune, and when to just write a better prompt.

Deployment and product surface

Why the same underlying model feels different inside every app that wraps it.

Teams and roles

How AI changes who does the work, starting with the Forward Deployed Engineer.

Spec-driven development

Writing the specification first, and where the evidence for it actually is.

Model behavior

Personality, consistency, and how models actually behave across prompts and runs.

Hallucination

When and why models make things up, how often it happens, and what actually reduces it.