What we kept the expensive model for, after moving the bulk of the work:
- ambiguous requirements, where the useful output is questions rather than code
- anything touching auth, permissions or money
- the final review pass on a large diff
Everything else, meaning the mechanical multi file edits, test writing and the "where does this happen" archaeology, went to the cheap tier and I have not missed it. The split is roughly planner and reviewer expensive, implementer cheap, and it survives contact with a big repo better than picking one model for everything.