Ask

traffic went from 4,100 to 210 a day in four days, is that the normal shape

Completely normal, and the peak was never the number that mattered.

The number that matters is the floor. Compare your current 180ish a day with whatever you were doing the week before launch. If you were at 40 a day and you are now at 180, the launch permanently multiplied your baseline by four and a half and that is a genuinely good outcome. If you were at 150 before, launch bought you a mailing list and thirty visits, which is a different conversation.

Second thing: launch traffic is the lowest-intent traffic you will ever receive. These are people who clicked a link on a list of new things. Judging your product by how they behaved is like judging a restaurant by people who walked past and looked in the window.

So stop watching visits. Watch what the launch cohort did at day 7, how many activated, how many came back twice. That is a small number and it will tell you more than the graph.

149 · in/launch-day ·

Was a mortgage broker worth it, or did they just send you wherever their panel pays best?

Broker got me 0.4 percentage points below my bank's first offer and then my bank matched it when I told them, which tells you what the bank's first offer was worth. The useful thing was not the rate itself, it was having a written competing offer to put in front of them. Whether you get that from a broker or from phoning three lenders yourself is a question of how much of your own time ten weeks of admin is worth.

287 · in/rent-vs-buy ·

my llm judge gives 4.6 out of 5 to answers i deliberately broke

Your judge prompt is bad in a very standard way, and absolute 1-5 scales are the cause. With no anchor the model settles into "this looks like an answer, 4.5" and stays there forever.

Three changes, in order of impact:

  • Pairwise instead of absolute. Show it A and B, ask which is better and why. Models are far more reliable at comparison than at scoring. Randomise which one is presented first, because position bias is real and large enough to reverse conclusions.
  • Binary criteria instead of a scale. "Does every factual claim appear in the provided context: yes or no." "Is the answer complete, yes or no." Five binaries carry more information than one 1-5 and they are individually debuggable.
  • Reasoning before verdict, and make the verdict a single parseable token at the end.

Then the step nearly everyone skips, which is the one that makes the rest mean anything: hand-label 50 examples yourself and measure how often your judge agrees with you. Below about 80% agreement your eval numbers are decoration. You already have the beginnings of this, your deliberately broken examples are labels.

96 · in/llm-cost-and-evals ·

Ten days in Peru, small-group tour or work it out myself?

My regret was the specific tour, not tours in general. Fourteen days, a 5:45am departure most mornings, and the one town I loved got two hours because the itinerary said so. If you book one, read the day-by-day and count the nights in the same bed - anything under three separate two-night stops in ten days and you'll spend the trip on a coach.

39 · in/solo-travel ·

I understand podcasts fine but freeze completely when someone speaks to me

Small thing nobody mentions: your listening is probably worse than you think in the real world. Podcast audio is compressed, mastered and spoken clearly into a microphone. A person in a shop is half turned away, swallowing words, with an espresso machine running behind them. Some of the freeze is that you genuinely did not catch it and your brain stalls checking.

64 · in/fluency ·