102You set temperature to 0 and send the same prompt twice, and the outputs differ. Why, and when does it matter?▼hardAnthropicOpenAIDatabricks◆ premiumTemperature 0 is not the same as deterministic, and the reason lives in the GPU kernels, not the sampler. What gets scored is naming the batch-invariance problem and knowing which fixes are real versus placebo.Open full answer →
123Why can't you evaluate an LLM application the way you evaluate a classifier?▼mediumOpenAIAnthropicDatabricks◆ premiumAccuracy against a held-out label is the wrong instrument for a system with no single right answer, an open failure space, and a metric that is itself a model. Here is what breaks, what survives, and what replaces it.Open full answer →