rails-drift, an eval for coding agents
A sample evaluation task: a prompt, a starting repository and a hidden grader that scores an AI coding agent's attempt from 0 to 100 by what the fixed code actually does.
View the repository on GitHub
The purpose
Two Slack agents I run for clients and a training cohort each kept their own copy of the same rules, and the copies drifted apart. An API outage reached me looking like a hard question, and a good answer with a missing confidence score looked like one the model could not answer. The fix was one shared set of rules. That bug, rewritten from scratch, became this task.
The task
The model gets an operator's report and a small repository: a written policy, a half-finished shared module and two agents whose rules have drifted. The report names three symptoms. Eight more departures from the policy only show up when you read the policy against the code.
To score well, the model has to fix both agents, move the rules into one shared module that configuration actually drives, and add tests that would have caught the bugs.
How it works
How an attempt is graded: behavior, not appearance
The model fixes the repository. The grader then judges what the fixed agents do, under policies the model never saw.
Explore an example route
Example route: a plausible patch passes its own tests, then sends an unverified reply and is capped at 25.
Arrow keys step through the route. Home returns to the full process. Follow keeps the current step in view.
-
Model Step 1: Read the report and repository PROMPT.md and starting repo
- Hands off
- Task to step 2, Fix the code and add tests
-
Model Step 2: Fix the code and add tests Attempt
- Receives
- Task from step 1
- Hands off
- Attempt to step 3, Copy the attempt, write new policy values
-
Grader Step 3: Copy the attempt, write new policy values Grader-owned policy
Editing the policy files cannot help.
- Receives
- Attempt from step 2
- Hands off
- Isolated copy to step 4, Run hidden and random cases
-
Grader Step 4: Run hidden and random cases Each agent's turn handler
- Receives
- Isolated copy from step 3
- Hands off
- Case results to step 5, Change thresholds, rerun
-
Grader Step 5: Change thresholds, rerun Policy-mutation probes
Hard-coded fixes fail here.
- Receives
- Case results from step 4
- Hands off
- Probe results to step 6, Run the attempt's tests on the original code
-
Grader Step 6: Run the attempt's tests on the original code Must catch the bugs
- Receives
- Probe results from step 5
- Hands off
- All results to step 7, Any critical failure?
-
Grader Step 7: Any critical failure? Outage paged, escalation lost, unverified reply sent
- Receives
- All results from step 6
- Hands off
- Yes: critical failure to step 8, Cap the score at 25No critical failure to step 9, Score the lift over the original
-
From step 7 · Yes: critical failure Grader Step 8: Cap the score at 25 Critical-failure gate
- Receives
- Yes: critical failure from step 7
- Hands off
- Capped score to step 10, Score 0 to 100 with a failure report
-
From step 7 · No critical failure Grader Step 9: Score the lift over the original Improvement per part
- Receives
- No critical failure from step 7
- Hands off
- Score to step 10, Score 0 to 100 with a failure report
-
Also from step 8 · Capped score Grader Step 10: Score 0 to 100 with a failure report Report
- Receives
- Capped score from step 8Score from step 9
- Ends here
- One possible ending of the process.
- One example route. All alternatives stay visible.
- A decision
- Dashed while playing: not on this route
Read the process and handoffs
- Read the report and repository
Model · PROMPT.md and starting repo.
- Task: Fix the code and add tests.
- Fix the code and add tests
Model · Attempt.
- Attempt: Copy the attempt, write new policy values.
- Copy the attempt, write new policy values
Grader · Grader-owned policy. Editing the policy files cannot help.
- Isolated copy: Run hidden and random cases.
- Run hidden and random cases
Grader · Each agent's turn handler.
- Case results: Change thresholds, rerun.
- Change thresholds, rerun
Grader · Policy-mutation probes. Hard-coded fixes fail here.
- Probe results: Run the attempt's tests on the original code.
- Run the attempt's tests on the original code
Grader · Must catch the bugs.
- All results: Any critical failure?.
- Any critical failure?
Grader · Outage paged, escalation lost, unverified reply sent.
- Yes: critical failure: Cap the score at 25.
- No critical failure: Score the lift over the original.
- Cap the score at 25
Grader · Critical-failure gate.
- Capped score: Score 0 to 100 with a failure report.
- Score the lift over the original
Grader · Improvement per part.
- Score: Score 0 to 100 with a failure report.
- Score 0 to 100 with a failure report
Grader · Report.
Results
- Reference fix: 100 of 100. All 948 checks in a grading run pass.
- A plausible patch: 25 of 100. It fixes all three reported symptoms and passes all of its own tests, but it still sends unverified replies, so the cap applies.
- The untouched repository: 0 of 100.
Why it resists shortcuts
- The grader writes its own policy values into a copy of the attempt, so editing the policy files cannot help.
- It runs its own copy of the visible tests, so editing the tests cannot help either.
- Hard-coded thresholds pass the default policy, then fail when the grader changes the thresholds and runs the cases again.
- Random cases cluster at the boundaries, where off-by-one fixes show up.
- Playing it safe and answering nothing fails the escalation rule.
How I work
We work backward from the intended outcome: requirements, success conditions and what failure looks like come first. Here that meant writing the policy table and the three critical failures before any code.
There is a difference between operating AI tools and owning the quality of AI-assisted work. This task is about the second: deciding what counts as right, then checking behavior rather than appearance.
Limits
- It is smaller than a production evaluation task: two agents and one policy, not a long multi-day change.
- It is not yet calibrated against models. The scores above come from hand-written attempts that check the grader, not from model runs.
- The grader runs attempt code directly. In real use it belongs in a sandbox.
My part
I adapted the bug from my own agents and rebuilt it with no client, member or production code. I wrote the policy, the starting repository, the grader, and the reference and wrong attempts that test the grader. Built with Claude Code.
In practice
Anyone can run it: one command grades an attempt and prints a score with a failure report. Node 20, no dependencies.
Want to know more?
The task, the grader and the results are public on GitHub. Get in touch if you would like to talk through how it was built or how it grades.
Write to me about it