Route by task shape, not by price, and move one workload at a time.
What has moved down a tier cleanly for me: schema-constrained extraction where the schema does the work, classification into a small label set, format rewriting, summarising a document that is already in context. Anything where the answer is mostly transcription under constraints.
What came back up: multi-step work with tools, where a cheap model's small planning errors compound across calls, and anything requiring it to notice that the document does not contain the answer. Abstention is consistently the first thing to go.
With 40M input tokens a month you are the ideal candidate for the move, but do it behind an eval and one workload at a time, so when quality dips you know which change caused it.
That is an assumption, and the announcement says otherwise: the stated cause was efficiency improvements across the training and inference stack, and the models on offer and their identifiers did not change. Serving cost falling over time is the normal shape of this industry, not evidence of a nerf.
Which is not the same as saying quality is guaranteed stable. Providers do change what sits behind a name. The correct response is the same either way: measure it on your own data and keep measuring, rather than inferring quality from the price in either direction.
Which is exactly why the used route dominates this budget. Both of those machines have been sold in enormous numbers for years, so the second-hand supply is deep and the prices are sane.
For investigation work specifically, grounding beats model choice most of the time. If the agent cannot pull the log slice and the trace itself, both models are guessing from whatever you pasted, and they converge on similar quality of guess.
Wire the tools first: log query, trace fetch, repo search that actually returns the right files. Then re-run your comparison. Every time I have done this the gap between models narrowed a lot, and the remaining gap was in how well each one used the tools rather than what it knew.
Check what the email looks like on a phone before you change anything else. Clinic owners read mail on an iPhone between patients. If your four lines render as nine with a signature block and an image, you're dead before the copy matters.
Also drop open tracking. It hurts deliverability slightly and it made you draw the wrong conclusion here, open rates have been noise since mail clients started prefetching images.
Optimized images are cached permanently after the first request, so this cannot be transformation volume unless your cache is being purged. I would look for a deploy hook or a revalidation call that is clearing the image cache on every deploy.
The router-then-load approach is more moving parts than I wanted, but the enum consolidation is free and I can do it this week. We definitely have three restart tools.
How did you label the expected tool for requests where a human would reasonably pick either of two? About a fifth of ours are genuinely ambiguous and I do not want to bake my own opinion in as ground truth.
Get three quotes and ask each installer for the production estimate in kWh/year plus the assumptions behind it. The spread between installers on the identical roof is embarrassing, and that number is what your whole payback calculation hangs on.
Buy a sample of a ten year old sheng from a similar region before you commit to storing anything for a decade. Then you know whether the destination is somewhere you want to go.
One test: can you finish the sentence "you can stop doing X now"? If yes, email. If the best you can manage is "we improved Y", it goes on the page and nowhere else. Ships that pass that test are rare and that is the point.
A changelog page is reference material, not a channel. Nobody wakes up and browses your release notes. Two things moved the needle for us:
a single line of in-app copy that only appears on the screen the change touches, dismissable, and only shown to accounts created before the ship date. Discovery of the feature it pointed at went from something like 4% of active accounts to 38% in a fortnight.
an email, but only for changes that alter something a person already does by hand. Roughly one entry in five qualifies. Everything else is noise and trains people to ignore you.
The page still earns its keep for a different reason: support links to it, and anyone evaluating you will read a year of entries to check the thing is alive. 11 views a month is a perfectly good number for that job.
Waiting periods are the thing to plan around. If you're taking cover because you expect to use it soon, read the pre-existing condition waiting period before you sign rather than at the point of claiming.
Look at the card. Nine times out of ten the twenty successful reviews were passes on a prompt with a giveaway in it, and the one honest test failed. Long intervals expose lazy cards.
Check what you can do once at upload rather than per question. Structure extraction, section splitting, summaries, an outline: all of it is per-document work that a lot of systems redo on every single request without anyone noticing.
I'd buy the cheap one and spend the difference on a second pot. Two 4-litre pots beat one 6-litre for how I actually cook, stock going in one, braise in the other, both on the hob at once. No warranty gives you that.
Test the card the way the exam tests you. If the exam asks you to name a structure from a description, your card should show the description and ask for the name. Cloze deletions pulled out of prose usually test neither direction properly.
Only if the percentage is the whole post. "Churn is 4%, here is the cancellation survey question that got us there" is not evasive, it is specific about something more interesting than revenue. People notice vagueness, not the absence of a dollar sign.
I did the opposite and do not regret it. Posted percentages and milestones - "first month over 20 customers", "churn down to 4%" - and never an absolute revenue figure while I was employed. It cost me approximately nothing in engagement, because the interesting part of a post is never the dollar amount, it is what you changed and what happened.
Also worth remembering the number outlives the moment. It gets screenshotted, it turns up in a search result, and in eighteen months you may be in a conversation about acquisition or a new job where you would rather choose what you disclose.
Went through this exact thing last year. New shoes made day one feel great and changed nothing about the underlying problem, which was that I ran every single run at the same slightly-too-hard pace. Slowing down about a minute per km did more than any purchase.
Ask what amount the policy is written for. Some owner's policies are written at purchase price with no inflation provision and others have an escalation feature. Twenty years from now that difference is effectively the whole policy.
Two months on one tin is worth mentioning too. Roasted oolongs fade, especially if the tin is opened daily and lives near a hob. Get a fresh sample of the same tea for your bottle test so you are only changing one thing.