Method: Measuring Translation Quality Drift Across Long Sessions (2026)
Supporting method page for Does AI Translation Get Worse Over a Long Meeting?. Published 2026-08-18.
Design
Each session is compared against itself. A block of eight consecutive exchanges from the opening of the session and eight from its close were scored independently, and the reported change is the mean of the per-session differences. Because the comparison is within a session, the speakers, topic, acoustics, accents and language pair are held constant by construction; nothing needs to be controlled for statistically because nothing varies between the two arms except position in the session.
Exchanges shorter than four words were skipped, as were any where the translation was byte-identical to the source, which indicates passthrough rather than translation. Sessions shorter than 20 minutes were excluded entirely.
Scoring
Each exchange was scored 0-5 for comprehension fidelity by three independent frontier LLM judges (GPT-4o-2024-11-20, Gemini 2.5 Flash, Claude Sonnet 4.6), and the value used is the arithmetic mean of the three. The rubric penalises wrong-word substitution, dropped entities or numbers, punctuation that changes sentence type, and language-script mixing; it does not penalise register or valid synonyms.
Complete panels only. An earlier run of this analysis averaged whichever judges happened to answer, and the GPT judge was intermittently rate-limited. Because the opening and closing blocks are scored sequentially within a session, a partial panel lands disproportionately on the closing block and manufactures an apparent decline. That artifact produced a spurious -0.312 before it was identified. The final run retries until all three judges respond and discards any exchange that cannot obtain a complete panel; all 2,128 scored exchanges carry three judges.
Headline result
133 sessions, 2,128 scored exchanges. Opening mean 2.5805, closing mean 2.5805, mean within-session change +0.000, bootstrap 95% confidence interval [-0.127, +0.128] over 5,000 resamples with a fixed seed. 62 sessions closed higher, 58 lower, 13 unchanged. Median session length 29.6 minutes, longest 121.7, spanning 41 language pairs.
By session length
| Band | Sessions | Opening | Close | Change |
|---|---|---|---|---|
| 20 to 30 minutes | 70 | 2.558 | 2.569 | +0.011 |
| 30 to 45 minutes | 40 | 2.555 | 2.594 | +0.039 |
| 45 to 70 minutes | 19 | 2.691 | 2.581 | -0.110 |
Per-language-pair breakdowns are not published. With 41 pairs across 133 sessions, several cells fall to single digits, where the numbers are dominated by noise and a rare corridor could narrow the field of possible sessions more than we are willing to publish.
Three-point shape check
A two-point comparison cannot distinguish steady quality from a mid-session dip that recovers. On a 60-session subset an additional block was scored at the midpoint: opening 2.6229, midpoint 2.3569, close 2.5868. Opening to close is -0.036; the midpoint sits 0.266 below the opening and the close sits 0.230 above the midpoint. Those excursions are well inside the between-session standard deviation of 0.742 and do not run in a consistent direction, so the shape is flat with noise rather than a slope.
Statistical power
The standard deviation of per-session change is 0.742 at n=133, giving a minimum detectable effect of approximately 0.18 points at 80% power. The result should be read as no decay larger than roughly 0.18 on a 0-5 scale, not as a demonstration that the change is exactly zero. Notably, the between-session spread of 0.742 is about twenty times that floor: which session you are in dominates how far into it you are.
The abrupt-failure measurement
Gemini 3.5 Live Translate was run on a 120-second Mandarin-to-English news clip. Across three runs, the last translation event arrived at 86,314 ms, 86,476 ms and 86,500 ms, while outbound audio events continued to 124,982-125,189 ms. The session therefore remained open and streaming while producing no further translation, silently dropping roughly 28% of the clip. The consistency of the cutoff across runs is why it is described as deterministic rather than as an intermittent fault.
Limitations
Long sessions only. Sessions under 20 minutes were excluded, so nothing here describes short sessions.
Endpoint blocks, not continuous sampling. Eight exchanges at each end, plus a midpoint block on the subset. A decay confined to a narrow window between sampled blocks would not be seen, though the three-point check makes that shape unlikely.
Judge resolution. Comprehension fidelity is a coarse instrument. A degradation that leaves meaning intact will not move it, which is precisely why the minimum detectable effect is reported rather than a bare claim of zero.
Independence. LiveLingo is a real-time translation product and this measurement was run by us. The headline result is a null that removes a differentiator we could otherwise have claimed about session length, and the abrupt-failure comparison is reported with the raw millisecond timings so it can be checked.