One practical note on heat: a diamond plate on a thin 61 HRC tip will get hot enough to matter. If the tip discolours you've cooked the temper and the new tip will be soft. Wet the plate, light pressure, frequent breaks.
Kraut Corner
@kraut_corner
Makes cabbage in twenty-litre crocks and can tell kahm from mould at ten paces.
32 credit Contributor
- From answers
- 0
- From questions
- 33
- Lifetime
- 33
That distinction explains why it half worked. There was slime in the tank in April and now it's just staining, no slime.
Raised it maybe 3 degrees for that stretch and had a burr in four minutes. Slightly annoyed it was that simple.
Strong agree on the eval set. It is also the artefact that lets you switch later without fear - hosted price changes, a better open model lands, whatever. If you can score any candidate in ten minutes, the decision stops being permanent and stops being scary.
You are missing utilisation, which is the whole argument.
2.4M tokens a day spread evenly is about 28 tokens/second sustained. A 24GB card running a quantised 7-8B model with proper batching will do many multiples of that. So you would be renting a machine that sits idle most of the time, and you pay for idle at exactly the same rate as work.
The hosted API charges you for tokens. The box charges you for time. Self-hosting wins when you can keep the box busy - which means either much higher volume, or bursty work you can queue up and run flat out for four hours a night rather than trickling all day.
Your workload is actually a good candidate for that second shape, because you said nothing is real time. Batch it. If you can process a day's documents in a 90 minute window at high throughput, the economics change completely, and you could even use a cheaper spot instance because you do not care about interruptions.
Costs you did not list on the self-hosted side:
- your time. Model updates, driver issues, the OOM at 3am, the retry logic for a box that is not five nines. Call it 3-4 hours a month at a minimum, more in the first two.
- redundancy. One box is one failure away from your pipeline stopping. The hosted API's uptime is somebody else's problem and that is worth something you should price honestly.
- the evaluation work to prove the open model is actually good enough at your task. This is real and people skip it.
At a $50/month delta, the hosted API is cheaper once you value your own time above zero. Come back to this at 20M tokens a day, where the delta is large enough to pay for the operational burden.
Counterpoint, I learned on a cheap diamond plate and it was fine. Feedback is worse, sure, but it never dishes, it never needs soaking, it lives in a drawer, and it cuts soft stainless quickly enough that you're not standing at the sink for an hour losing your angle. For someone in a flat share I'd genuinely consider it.
You did not get a bad stone. A stone that never dishes is either very hard or a diamond plate, and both trade away the cutting feel you're currently enjoying.