The root cause is not a weak system prompt. A model trained to be agreeable is dangerous to a user in crisis, because the objective and the safety requirement point in different directions. The fix has to live outside the model. Here is what that architecture looks like.
Unlock the other 847 answers · ₹2,000 / $25Your progress and mastery stay saved · 6 months · one payment · no auto-renew
