All projects

Published sample task · October 2026

rails-drift, an eval for coding agents

A sample evaluation task: a prompt, a starting repository and a hidden grader that scores an AI coding agent's attempt from 0 to 100 by what the fixed code actually does.

View the repository on GitHub
Two black paper rails run across a warm paper set toward an arch, with a red gauge block wedged between them and a blue thread tied to a blank paper card.

The purpose

Two Slack agents I run for clients and a training cohort each kept their own copy of the same rules, and the copies drifted apart. An API outage reached me looking like a hard question, and a good answer with a missing confidence score looked like one the model could not answer. The fix was one shared set of rules. That bug, rewritten from scratch, became this task.

The task

The model gets an operator's report and a small repository: a written policy, a half-finished shared module and two agents whose rules have drifted. The report names three symptoms. Eight more departures from the policy only show up when you read the policy against the code.

To score well, the model has to fix both agents, move the rules into one shared module that configuration actually drives, and add tests that would have caught the bugs.

How it works

How an attempt is graded: behavior, not appearance

The model fixes the repository. The grader then judges what the fixed agents do, under policies the model never saw.

  1. Model Step 1: Read the report and repository PROMPT.md and starting repo
    Hands off
    Task to step 2, Fix the code and add tests
  2. Model Step 2: Fix the code and add tests Attempt
    Receives
    Task from step 1
    Hands off
    Attempt to step 3, Copy the attempt, write new policy values
  3. Grader Step 3: Copy the attempt, write new policy values Grader-owned policy

    Editing the policy files cannot help.

    Receives
    Attempt from step 2
    Hands off
    Isolated copy to step 4, Run hidden and random cases
  4. Grader Step 4: Run hidden and random cases Each agent's turn handler
    Receives
    Isolated copy from step 3
    Hands off
    Case results to step 5, Change thresholds, rerun
  5. Grader Step 5: Change thresholds, rerun Policy-mutation probes

    Hard-coded fixes fail here.

    Receives
    Case results from step 4
    Hands off
    Probe results to step 6, Run the attempt's tests on the original code
  6. Grader Step 6: Run the attempt's tests on the original code Must catch the bugs
    Receives
    Probe results from step 5
    Hands off
    All results to step 7, Any critical failure?
  7. Grader Step 7: Any critical failure? Outage paged, escalation lost, unverified reply sent
    Receives
    All results from step 6
    Hands off
    Yes: critical failure to step 8, Cap the score at 25No critical failure to step 9, Score the lift over the original
  8. From step 7 · Yes: critical failure Grader Step 8: Cap the score at 25 Critical-failure gate
    Receives
    Yes: critical failure from step 7
    Hands off
    Capped score to step 10, Score 0 to 100 with a failure report
  9. From step 7 · No critical failure Grader Step 9: Score the lift over the original Improvement per part
    Receives
    No critical failure from step 7
    Hands off
    Score to step 10, Score 0 to 100 with a failure report
  10. Also from step 8 · Capped score Grader Step 10: Score 0 to 100 with a failure report Report
    Receives
    Capped score from step 8Score from step 9
    Ends here
    One possible ending of the process.
  • One example route. All alternatives stay visible.
  • A decision
  • Dashed while playing: not on this route
Read the process and handoffs
  1. Read the report and repository

    Model · PROMPT.md and starting repo.

    • Task: Fix the code and add tests.
  2. Fix the code and add tests

    Model · Attempt.

    • Attempt: Copy the attempt, write new policy values.
  3. Copy the attempt, write new policy values

    Grader · Grader-owned policy. Editing the policy files cannot help.

    • Isolated copy: Run hidden and random cases.
  4. Run hidden and random cases

    Grader · Each agent's turn handler.

    • Case results: Change thresholds, rerun.
  5. Change thresholds, rerun

    Grader · Policy-mutation probes. Hard-coded fixes fail here.

    • Probe results: Run the attempt's tests on the original code.
  6. Run the attempt's tests on the original code

    Grader · Must catch the bugs.

    • All results: Any critical failure?.
  7. Any critical failure?

    Grader · Outage paged, escalation lost, unverified reply sent.

    • Yes: critical failure: Cap the score at 25.
    • No critical failure: Score the lift over the original.
  8. Cap the score at 25

    Grader · Critical-failure gate.

    • Capped score: Score 0 to 100 with a failure report.
  9. Score the lift over the original

    Grader · Improvement per part.

    • Score: Score 0 to 100 with a failure report.
  10. Score 0 to 100 with a failure report

    Grader · Report.

    Results

    • Reference fix: 100 of 100. All 948 checks in a grading run pass.
    • A plausible patch: 25 of 100. It fixes all three reported symptoms and passes all of its own tests, but it still sends unverified replies, so the cap applies.
    • The untouched repository: 0 of 100.

    Why it resists shortcuts

    • The grader writes its own policy values into a copy of the attempt, so editing the policy files cannot help.
    • It runs its own copy of the visible tests, so editing the tests cannot help either.
    • Hard-coded thresholds pass the default policy, then fail when the grader changes the thresholds and runs the cases again.
    • Random cases cluster at the boundaries, where off-by-one fixes show up.
    • Playing it safe and answering nothing fails the escalation rule.

    How I work

    We work backward from the intended outcome: requirements, success conditions and what failure looks like come first. Here that meant writing the policy table and the three critical failures before any code.

    There is a difference between operating AI tools and owning the quality of AI-assisted work. This task is about the second: deciding what counts as right, then checking behavior rather than appearance.

    Limits

    • It is smaller than a production evaluation task: two agents and one policy, not a long multi-day change.
    • It is not yet calibrated against models. The scores above come from hand-written attempts that check the grader, not from model runs.
    • The grader runs attempt code directly. In real use it belongs in a sandbox.

    My part

    I adapted the bug from my own agents and rebuilt it with no client, member or production code. I wrote the policy, the starting repository, the grader, and the reference and wrong attempts that test the grader. Built with Claude Code.

    In practice

    Anyone can run it: one command grades an attempt and prints a score with a failure report. Node 20, no dependencies.

    Want to know more?

    The task, the grader and the results are public on GitHub. Get in touch if you would like to talk through how it was built or how it grades.

    Write to me about it

    Ask Fred

    AI guide to this site

    1. Ask Fred · AI

      Hi. I can point you to Fred's projects and notes, and help you get in touch. What would you like to know?

    AI can make mistakes. Keep private details out.

    Get in touch