Copied to clipboard!
Module 9 Foundations intermediate 36 min

Building an evaluation harness

What you'll be able to do

  • Set up a cheap evaluation with reference outputs and an A/B comparison in a Google Sheet
  • Decide whether a prompt change actually improved things, with evidence instead of a hunch
  • Make "good" measurable by scoring two or three criteria that matter, not one fuzzy sense of better

Build a Cheap Evaluation Harness in a Spreadsheet

"It feels better" is not proof. When you change a prompt, you need a way to tell real improvement from luck. This exercise builds the cheapest evaluation that works: a few reference inputs, an answer key for what good looks like, and an A/B scorecard in a plain Google Sheet. You change one thing at a time and let the evidence decide. You stay the judge of what "good" means; the harness just keeps you honest.

Recommended tool

claude

Claude is good at proposing concrete, distinct scoring criteria and laying out a clean spreadsheet structure, which is exactly the scaffolding you need before you start scoring outputs by hand.

Bring to the exercise

  • A prompt you run often and have been tempted to tweak
  • Three to five real inputs you can test it with (with names removed if sensitive)
  • A rough sense of what a great output looks like for each input
Step 1

Pick Criteria That Actually Matter

Replace one fuzzy "is it better?" with two or three criteria you can score, so improvement becomes measurable.

What to substitute before pasting

  • [DESCRIBE THE TASK] What the prompt is for, in a sentence or two.
Help me build a cheap evaluation for a prompt I use a lot.

The task this prompt does:
[DESCRIBE THE TASK]

Propose exactly three scoring criteria that actually matter for this task (for example: correct facts, right tone, followed the format). For each, give me a one-line definition and a simple way to score it (pass or fail, or 1 to 3). Do not give me a single vague "quality" score.
What good output looks like
  • Exactly three criteria, each distinct and tied to this task.
  • Each criterion has a definition and a simple scoring method.
  • No catch-all "overall quality" score doing all the work.

Verify before using AI's output

Step 2

Build the Answer Key

Reference outputs are your answer key. A handful of inputs with the "this is what good looks like" output gives you a baseline.

What to substitute before pasting

  • [PASTE OR DESCRIBE 3-5 INPUTS] Three to five genuine inputs, varied enough to stress the prompt, names removed if sensitive.
Here are real inputs I can test with:
[PASTE OR DESCRIBE 3-5 INPUTS]

For each input, help me write a short reference output: a description of what a strong answer would include, so I have an answer key to compare against. Do not just generate slick answers; describe the qualities a winning answer must have for this input.
What good output looks like
  • Each input has a reference describing what a strong answer must contain.
  • The references describe qualities, not just a single perfect wording.
  • The set covers a real range, not three near-identical cases.

Verify before using AI's output

Step 3

Lay Out the A/B Scorecard

Build a scorecard in a plain sheet so you can run the old prompt and the new one on the same inputs and compare.

Lay out a Google Sheet structure for A/B testing two versions of my prompt against these inputs and my three criteria.

Give me:
- The exact columns (input, output A, output B, score per criterion for each, who won, why)
- One example row filled in so I can see how it is used
- A one-line reminder of the rule that I change only ONE thing between version A and version B

Keep it simple enough that I will actually run it every time I change the prompt.
What good output looks like
  • Columns cover input, both outputs, per-criterion scores, a winner, and a reason.
  • The example row shows the sheet in use.
  • It states the one-change-at-a-time rule plainly.

Verify before using AI's output

Step 4

Decide How Much Evidence Is Enough

Set the bar for trusting the result, sized to the stakes, so you do not over- or under-test.

What to substitute before pasting

  • [ONE LINE ON WHAT IT COSTS IF THIS PROMPT IS WRONG] The real cost of a bad output here, so the test size matches the stakes.
For a task at this stakes level:
[ONE LINE ON WHAT IT COSTS IF THIS PROMPT IS WRONG]

Tell me how many test inputs is "enough to trust" a decision between two prompt versions, and why. Then tell me what would make you want more tests, and what a result is NOT allowed to tell me (the limits of a small eval like this).
What good output looks like
  • Gives a concrete number of inputs tied to the stakes you described.
  • Names what would justify testing more.
  • Is honest about the limits of a small, cheap eval.

Verify before using AI's output