Release practice · 9 December 2025
Why lab scores drift after release
A gate that was green on Thursday can be a support tag by Monday. The drift is rarely magic. In Lab Protocol Design we keep finding the same three quiet causes: build flavour, thermal history, and fixture version.
Build flavour is the most embarrassing. A weekend lab run used a debug overlay or a different feature flag. The trace looked fine. Production did not. The exclusion-rules exercise exists so you write down which flavour is allowed in a published score. If you cannot name the flavour, you do not publish.
Thermal history is the one Bangkok makes obvious. A phone that sat in air-conditioning is not the phone that sat in a bag on the BTS. Module 03 of Field Benchmark Mastery runs the same scroll after a 12-minute soak. Teams who skip the soak keep a lab delta that never appears in the field. Our last Bench Lead cohort’s median lab-to-field delta was 11 ms after they adopted the soak; before that, several teams were quoting gaps they could not reproduce.
Fixture version is the one people forget to write down. TH-BKK-COMMUTE-12 is not TH-BKK-COMMUTE-14. Carriers change. If your protocol does not pin the fixture, your “regression” may be a new paging pattern. We retire fixtures in public on the lab page rather than silently editing them.
Lab-versus-field is not a reason to abandon the lab. It is a reason to treat the lab as a hypothesis generator. Field traces remain the argument. That sentence is on the wall in Wattana for a reason, and it is the opening pull quote on the home page because we still watch teams forget it under deadline.