Product notes / 2026
Making AI skill visible.
Fluency Bench is a work-sample assessment for how people judge, verify, and direct AI.

The problem
“I use AI every day” tells you very little about someone’s judgment. Can they catch a confident mistake? Turn a vague request into usable instructions? Notice when the workflow needs a person?
I’m building Fluency Bench around work you can inspect: the output someone chooses, the evidence they check, and the instructions they write.
Three things to watch
Judge the output
A polished summary is not necessarily the right one. The first round asks which output serves the actual task.
Check the claim
A recommendation arrives with seeded mistakes and plausible distractions. The task is to verify it against the sources.
Test the instructions
The workflow faces unseen cases. A prompt that reads well still has to produce the right behavior.
The decisions that matter
- Make the rubric inspectable.
- The public methodology publishes scoring anchors. That gives someone a concrete standard to question or practice against.
- Separate checking from judging.
- The methodology distinguishes deterministic checks from model-judged criteria and exposes large disagreements between judges. A single number should not hide how it was produced.
- Test beyond the examples.
- The published design uses held-out cases and reports the gap from the visible set. The question is whether the instructions generalize.
Try it, then look under the hood.
The product and its public rubric are the best starting points for a conversation about the work.