On contesting: know what it costs before you decide. The processor charges a fee when a dispute is raised, and since a change a while back several of them also charge a second fee when you choose to fight it, refunded only if you win. So a $12 dispute can turn into a meaningfully larger loss for the satisfaction of being right.
With logs this good I would still submit, because the evidence is strong and card networks do reverse these, but do it with your eyes open about the arithmetic. Check your own dashboard for the exact fees rather than trusting anyone's numbers including mine, they have moved more than once recently.
Pair it rather than replacing. Video for "how do I do the thing", one page for credentials, renewal dates and who to call in what order.
Video is useless at 2am. Nobody scrubs through a recording looking for the hosting login while panicking. The one-pager in their inbox is what gets found under stress, and the video is what stops the calm-daytime questions.
Retainer, but sell availability rather than hours.
The moment a retainer is denominated in hours, the client starts counting them and you start justifying them, and at the end of the month somebody feels cheated regardless of what happened. Denominate it in response: "critical issues answered within two business hours, everything else within one working day, this many changes per month included."
That gets signed more easily in my experience because it maps to what they are actually anxious about, which is not hours. It is being ignored when something breaks.
I price it as a percentage of the build cost per month and adjust after two quarters of real data. That is my rule of thumb, not an industry standard, and the honest input is: what would it cost me to be reliably interruptible for this client, plus what I lose by holding that capacity.
The fudge factor is not a failure of method. It is the price of the option you are selling them.
First, require consecutive failures. A single 60 second check firing on the first bad response is a blip detector, not an outage detector. Three consecutive failures means you find out about three minutes late and you stop hearing about every deploy, every certificate renewal, every provider hiccup. Nearly all of my noise was in that window.
Second, split the alert channels by what you would actually do. Down and staying down goes to push. Everything else - error rate up, queue backing up, disk at 80% - goes to an email that you read with coffee. If you would not get out of bed and open a laptop for it, it is not a page, it is a note.
Disagreeing with the single-payer approach slightly. It concentrates all the fraud risk, all the card-blocked-abroad risk and all the admin on one person, and if his card gets frozen on day two the whole system stops. Two payers alternating, both logging into the same app, has been more robust for us and it barely costs anything if both cards are reasonable on fees.
Docs deflect, but only if the answer is findable from inside the product at the moment of confusion. A help centre nobody links to is a blog. I put a small link next to each confusing control that goes straight to the relevant paragraph, and support volume for those features dropped noticeably. Not zero. Noticeably.
I bought with almost nothing left after the deposit and the boiler failed in month three, which turned a cheap month into a horrible year. Whatever the gap is, do not spend it in advance. Keep three to six months of the full payment plus a repair fund and the whole thing stops being frightening.
Bedding did more for me than any of the airflow stuff. A linen flat sheet and no duvet at all in summer, plus a pillow that does not hold heat. Cotton percale is the cheap version if linen is out of budget. Anything with polyester in it in July is a mistake.
Model it before you commit, and get the current numbers yourself rather than from a thread, because the fee structures and the cost of bidding have both changed more than once since I started and anything I quote you will be out of date. What matters is the arithmetic: cost per proposal, proposals per win, platform cut of the win. Mine worked out to a meaningful percentage of gross once bidding costs were in, which was fine at the rates I ended up at and would have been ruinous at the rates I started at.
Ours was 31 percent and comfortable, and the thing that made it comfortable was fixing the rate for long enough to see the childcare years out. Worth asking a broker what the fix costs you against the flexibility you give up, because that trade off is very personal and I got it wrong the first time.
Eight concurrent users on one card is exactly the case vLLM is built for, so yes. The mechanism is continuous batching: instead of finishing one request and starting the next, it keeps a running batch and slots new sequences in as others finish, so the GPU stays busy. Single-user speed will be roughly what you have now, maybe a little better; aggregate throughput under concurrency is several times higher in my testing, and more importantly the tail latency stops being awful for whoever arrives last. The cost is that it is a server, not an app - you pin a model, you configure memory, and it is much less forgiving than Ollama about changing your mind.
some defence of the percentage, though. it is useless month to month at your size, but smoothed over a quarter it is the first thing that tells you a stall is structural rather than seasonal. i would not throw it out, i would stop looking at it monthly and put it on a quarterly review instead.
Fair, but the fix is not necessarily 'build sooner'. It is 'stop delivering within an hour by hand and start delivering at a fixed time'. If a 9am daily batch keeps everyone happy you have learned something real about the tolerance for latency and still not written a line.
I put a Japanese knife through a carbide pull-through exactly once, before I knew any better, and took two visible chips out of the edge along with a set of scratches. Grinding those out cost me a lot of blade height and an afternoon on a coarse stone, and the knife has a slightly different profile to this day. It was a present. If you take nothing else from this thread: hard thin steel and carbide scrapers do not mix.
Paid on-call is worth more than it costs for a second reason: it makes the expectation explicit rather than you begging someone by text on their day off.
The "then I start thinking about work" part is doing more work in this story than it looks. I keep a notepad by the bed and write the thought down in one line, no elaboration. It sounds like a productivity gimmick and it is not - it just stops the loop of rehearsing something so you do not forget it.
Take it if it generalises. Charge double if it does not. Never build it for free to keep someone happy - that customer never becomes happier, they become a habit.
Two disputes in six years. Both took several weeks of back and forth, both ended in a partial outcome rather than anyone winning outright, and in both cases my hourly rate for the hours I spent on the dispute was terrible. That is worth knowing in advance: at 2400 it is clearly worth fighting, and below a few hundred you are usually better off writing it off and never working with them again.