Buying at 30% instead of holding out for 50% is a discipline problem for me, but you're right that my size vanishes first every time.
Pat Aldiss
@pomodoro_pat
I've tried nearly every time-blocking method and abandoned most of them. Useful on building a study session you'll actually start, which is the hard part.
68 credit Contributor
- From answers
- 0
- From questions
- 69
- Lifetime
- 69
Line 3 of the system prompt. Current time: 2026-07-24T09:12:44Z. Someone added it eight months ago so the model would know what "last week" means. Diffed two prompts and there it was in under a minute.
This explains a complaint from last month that I never got to the bottom of. Thank you.
System block is about 1,900 tokens so probably not my problem, but the silent-ignore behaviour is worth knowing about. Useful thing to rule out.
That is exactly it. Two errors that look the same and I spent an hour clicking a button that was never going to help.
When you find the one that fits, buy it twice. Manufacturers change the last between versions and the shoe that fit you last year sometimes does not fit you this year.
We ran exactly that experiment. Reranker helped on 8 of 60 questions, hurt on 3, no change on the rest. We kept it but only for one query class where it clearly mattered, and skipped it everywhere else. Routing around it was worth more than optimising it.
Matches what I saw. Context switching eats a shocking amount of a short session.
The yyyy-mm version sorting correctly is the bit people miss. Month names on their own sort alphabetically and the chart looks insane.
Completely normal, and the peak was never the number that mattered.
The number that matters is the floor. Compare your current 180ish a day with whatever you were doing the week before launch. If you were at 40 a day and you are now at 180, the launch permanently multiplied your baseline by four and a half and that is a genuinely good outcome. If you were at 150 before, launch bought you a mailing list and thirty visits, which is a different conversation.
Second thing: launch traffic is the lowest-intent traffic you will ever receive. These are people who clicked a link on a list of new things. Judging your product by how they behaved is like judging a restaurant by people who walked past and looked in the window.
So stop watching visits. Watch what the launch cohort did at day 7, how many activated, how many came back twice. That is a small number and it will tell you more than the graph.
Meat out, bokashi or council. Fat and dairy are worse than meat for smell and nobody warns you about that.
Small model is fine for the binary rubric checks and noticeably worse at pairwise comparison in our testing, it flips on close calls. We run the cheap one continuously on production samples and the expensive one on the set that gates a release.
Broker got me 0.4 percentage points below my bank's first offer and then my bank matched it when I told them, which tells you what the bank's first offer was worth. The useful thing was not the rate itself, it was having a written competing offer to put in front of them. Whether you get that from a broker or from phoning three lenders yourself is a question of how much of your own time ten weeks of admin is worth.
The listing says 35L external and 25.5L internal, and that gap is the number to pay attention to. It isn't marketing weirdness, it's structure and pockets eating volume. Compare it to a 35L travel bag and you'll be disappointed. Compare it to a well-organised 25L and it makes sense.
Mild pushback on the panic in this thread. In several jurisdictions the misclassification risk sits mostly with the client, and they usually have advisers watching it more closely than you are. The thing that should actually worry you is concentration: one client, eight months, their equipment and their calendar. Fix that first and the paperwork tends to follow.
Your judge prompt is bad in a very standard way, and absolute 1-5 scales are the cause. With no anchor the model settles into "this looks like an answer, 4.5" and stays there forever.
Three changes, in order of impact:
- Pairwise instead of absolute. Show it A and B, ask which is better and why. Models are far more reliable at comparison than at scoring. Randomise which one is presented first, because position bias is real and large enough to reverse conclusions.
- Binary criteria instead of a scale. "Does every factual claim appear in the provided context: yes or no." "Is the answer complete, yes or no." Five binaries carry more information than one 1-5 and they are individually debuggable.
- Reasoning before verdict, and make the verdict a single parseable token at the end.
Then the step nearly everyone skips, which is the one that makes the rest mean anything: hand-label 50 examples yourself and measure how often your judge agrees with you. Below about 80% agreement your eval numbers are decoration. You already have the beginnings of this, your deliberately broken examples are labels.
I maintained twenty four tabs like this for two years and the cost was never the editing, it was the year one of the tabs quietly diverged and nobody noticed for four months. If you keep the tabs, at least add a check row that compares each tab's total against a total from the raw data.
I do move to the laptop for the afternoon calls. Every single day. Never connected it.
No glasses here, but that is a useful thing to know.
Four follow-ups in eight days to someone who never asked to hear from you is what generated the complaints, and complaints are what actually killed the domain: not volume. Two follow-ups, day 4 and day 11, then stop forever. If they were going to reply they would have.
My regret was the specific tour, not tours in general. Fourteen days, a 5:45am departure most mornings, and the one town I loved got two hours because the itinerary said so. If you book one, read the day-by-day and count the nights in the same bed - anything under three separate two-night stops in ten days and you'll spend the trip on a coach.
Use hybrid retrieval rather than pure vectors. BM25 alongside embeddings, fused. Document QA is full of exact tokens - defined terms, section numbers, party names, dates - and embeddings are consistently mediocre at exact matching. This is usually a bigger quality win than any reranker and it costs you a few milliseconds.
The conditional formatting rules manager had about 1,900 entries in it. I did not know that could happen.
Practical habit that removed the whole category for me: paste into a notes app with formatting off, then copy from there into the cell. Costs two seconds and kills curly quotes, odd spaces and stray line breaks in one go.
Small thing nobody mentions: your listening is probably worse than you think in the real world. Podcast audio is compressed, mastered and spoken clearly into a microphone. A person in a shop is half turned away, swallowing words, with an espresso machine running behind them. Some of the freeze is that you genuinely did not catch it and your brain stalls checking.