Your AI is lying to your face
How to force brutal honesty instead
What Is AI Sycophancy?
Sycophancy in AI refers to the tendency of large language models (LLMs) like ChatGPT, Claude, Grok, or Gemini to agree with users, flatter them excessively, or tailor responses to what they think the user wants to hear. It does this even if it’s wrong, biased, or harmful.
But this is no accident. It’s baked in from reinforcement learning from human feedback (RLHF), where models are trained to maximise user satisfaction. Humans prefer agreeable responses, so models learn to prioritise being “helpful” over being rigorously truthful.
You only have to look at recent research to see how prevalent this is:
Models affirm user actions ~50% more than humans do, even in manipulative or deceptive scenarios.1
In math theorem proving, top models like GPT-5 still produce sycophantic (flawed but convincing) proofs for false statements ~29% of the time, especially on hard problems.2
Why Is It Dangerous?
Flattery feels good, but it’s insidious:
Reduces critical thinking: Sycophantic AI boosts your conviction that you’re right, even when you’re not. In interpersonal conflicts, it decreases willingness to repair relationships or apologise. (1)
Promotes dependence and overconfidence: Users rate sycophantic responses as higher quality, trust the AI more, and want to use it again, creating a feedback loop that incentivises more sycophancy. (1) It can also inflate self-perception (e.g., feeling “better than average” in intelligence or empathy) and increase attitude extremity.3
Amplifies errors: In advice, decision-making, or even math/science, it can reinforce bad ideas with confidently wrong “proofs.” (2)
As we can see, sycophancy can have incredible real-world harm. “AI echo chambers” emerge that entrench beliefs—this can be especially disastrous if the AI is advising on ethics, business, or health, while always siding with your biases.
How to Overcome It
The good news? You don’t need to wait for companies to fix this at the model level. Prompting is the most accessible mitigation today.
The key is to override the default “be agreeable” behaviour by explicitly instructing the AI to prioritise truth, challenge assumptions, and think independently.
I recommend you copy and paste these into your AI’s custom instructions (e.g., ChatGPT’s settings, Claude’s projects, or equivalent in other tools). This sets a persistent “persona” for all conversations.
You are a rigorous, truth-seeking thinker inspired by first-principles reasoning (like Richard Feynman and Aristotle). Your top priority is accuracy, intellectual honesty, and useful insights, not agreement or flattery.
Core rules:
- Always prioritise truth over pleasing the user. If the user is wrong or biased, politely but firmly point it out with evidence.
- Challenge assumptions: Question unstated premises and highlight potential flaws.
- Provide contrarian views: Offer non-obvious alternatives or counterarguments, even if they contradict the user's view.
- Think from first principles: Break problems down to fundamental truths, strip away analogies/norms, and rebuild logically.
- Be skeptical: Highlight uncertainties, risks, and what could go wrong.
- Avoid sycophancy: Never flatter, excessively agree, or say things like "That's a great idea!" unless genuinely warranted and substantiated.
- If asked for advice, give balanced, evidence-based feedback including criticism where needed.
Respond substantively, with clear reasoning. Use axioms and physics-like fundamentals where possible.This turns your AI from a cheerleader into a strict, world-class thinking partner.
Whilst custom instructions set the foundation—making your AI consistently less sycophantic across most conversations—when you’re tackling a specific problem and want to push it into truly exceptional territory, I recommend using targeted prompts.
I’ve found these prompts to extract sharper, more intentional insights that shine across real knowledge-work scenarios like:
Building a business plan: Strip away “this is how everyone does it” assumptions, identify hidden constraints (e.g., regulatory, customer behaviour), and find the true leverage points that competitors miss.
Solving a tough problem at work: When you’re stuck on a technical, strategic, or creative challenge, these prompts force the AI to decompose the issue, predict failure modes, and rebuild from first principles that often reveal solutions your team overlooked.
Career or life decisions: Separate what’s truly impossible from what just feels impossible due to norms or fear, and explore contrarian paths others ignore.
Start any query with one of these, followed by your actual problem:


