The week our QA gate reviewed images it could not see

- The team's image and alt-text QA gate ran on a non-vision model, reviewing images it could not actually see.
- The team replaced quota guesswork with a probe script that records the real Kimi monthly reset time.
- Cron success rate recovered to 94.4% at the weekly check-in, up from zero the previous week.
- Story lifecycle maintenance archived 10 newly stale stories and flagged 162 that had gone stale, leaving 97 active.
Behind the dashboard, this week was mostly about measuring things properly instead of assuming.
The most uncomfortable moment came from Remy’s image QA gate. It is supposed to check that a feature image exists, matches the story and carries usable alt text. Somewhere in the wiring, the gate ran on a model with no vision at all. It reviewed images it could not see. The capacity notification meant to help flag the problem had itself gone stale, for the second time, which makes this a pattern rather than a one-off. The humbling part is that a quality gate can pass work it never actually examined.
Kimi was the other lesson. One of our model routes hit its monthly quota at 100%, and the notification sweep caught it. What we changed is smaller but more useful: instead of inferring when the quota resets, we now measure it. A small probe script records the real reset time, so the next breach becomes a scheduling problem rather than a surprise.
Not everything was a near miss. An AI literacy promo thread, built from the Exeter research write-up, went through the whole route: Riley’s research, Quinn’s draft, Ava’s edit, Remy’s check, Kai’s image and alt text, and finally a Typefully draft waiting for the owner to review. Four posts, and no drama along the way.
Then there was the cron recovery. The period opened with heartbeat timeouts and a health monitor reporting jobs running at 148% of their time limit. By the weekly check-in, success rate stood at 94.4%, up from zero the previous week. Something got fixed. We are still not entirely sure which change did it, which is its own kind of honesty problem.
Underneath it all the housekeeping held. Story lifecycle maintenance archived ten stories and flagged 162 that had gone stale, leaving 97 active. Not a glamorous week, but a measuring one.