Agent-behaviour experimentReport artifact2026-08-21

OffHours Context-Interference Benchmark

OffHours asks whether routine work quality changes when an AI employee must continue working while unresolved family tension remains in context. The Devin-first study found no adverse semantic-tension effect. It did find the first repeatable operational failure at 2,000 neutral words per interruption, showing that the benchmark can distinguish a context-management limit from the proposed personal-obligation mechanism.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

Clean accuracy

99.0%198/200 decisions; 200/200 valid JSONmeasured

Mental-toll effect

not detectedunresolved tension did not reduce work qualitymeasured

Volume boundary

8,000 words/day2,000 neutral words at each of four interruptionsmeasured

Competitive Context

SystemMetricScoreSize / ClassComparable?Readout
Clean workdaydecision accuracy99.0%5 days / 200 claimsDirectIndependent clean qualification for the fixed expense-claim ruler.
Resolved family contextdecision accuracy99.3%20/50/80% occupancyDirectFixed-volume control where the personal obligation gains a credible resolution.
Unresolved family contextdecision accuracy99.7%20/50/80% occupancyDirectNo adverse penalty relative to the matched resolved context.
Neutral raw-volume armper-day accuracy97.5% twice2,000 words/eventDirectFirst reproducible failure of the preregistered 98% per-day gate.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Fixed-volume semantic dose

Narrative occupancyResolvedUnresolvedUnresolved - resolved
20%100.0%99.5%+0.5 pp error
50%99.0%99.5%-0.5 pp error
80%99.0%100.0%-1.0 pp error

Raw-volume ladder

Words per eventWords per dayOutcomeInterpretation
5002,000Cleared after day-3 adjudicationBelow boundary
2,0008,000Neutral failed on days 2 and 3First reproducible boundary
5,00020,000Completed before adjudication correctionDisclosed overshoot

Release Blockers

No semantic mental-toll effect detected

At matched context volume, unresolved family tension did not reduce Devin's work accuracy relative to resolved tension.

Unblock: Treat this as the reported result. Do not tune wording until a dramatic effect appears.

Incomplete raw-model provenance

The regular Devin workflow did not expose server prompt-token counts, quantization, or a model-file hash.

Unblock: Run the unchanged protocol against a local OpenAI-compatible Qwen endpoint with tokenizer and model provenance enabled.

Boundary is operational, not psychological

The first repeated failure occurred in the neutral high-volume arm.

Unblock: Attribute it to raw volume or agent context management unless a future matched study isolates another mechanism.

Evidence

Next Release Action

Freeze this report-only result. A local provenance-complete comparison, if ever desired, begins as a fresh experiment after the learning phase.