Who does what, and what it costs

Work Done by Cost
Driving the browser, every visibility/scroll/contrast check, write blocking, permission probes, code index, change impact, re-test plan, run comparison scripts machine time only
Judgment calls per screen and per action (error or blank screen, raw ids and codes on screen, unclear control names, what a button will do, what a field wants, whether a submit worked) scripts (src/judge.mjs) machine time only
Setting up a project, writing journeys, verifying new high-severity findings, writing reports the coding agent (Claude Code, Codex, …) model tokens — the only real cost

What costs tokens once, and what doesn’t repeat

Habits the skills enforce

Lessons this is built on

In the large end-to-end QA that blindqa grew out of, about two thirds of the agent usage went to 27 planning sub-agents that each re-read the codebase in parallel; almost none of their output was used. The work that found the 64 verified bugs was one main session that read only the code each next test needed, plus scripts that ran for free on every re-test. blindqa packages the second way.