Language Models Don't Always Say What They Think¶
Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. Advances in Neural Information Processing Systems 36.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Enthymeme
- In AI reasoning, models produce arguments with hidden premises, and faithfulness research asks whether the stated reasoning is the actual reasoning.
This sourceShows that a model's stated chain-of-thought can systematically misrepresent its actual reasoning — supporting that AI reasoning carries hidden premises and that faithfulness asks whether the stated reasoning is the actual reasoning.
- In AI reasoning, models produce arguments with hidden premises, and faithfulness research asks whether the stated reasoning is the actual reasoning.
- Informal Fallacy
- In an AI deployment, a language model is asked whether a policy is sound and returns a fluent, confident argument that turns on the claim "every credible economist agrees," a fabricated-but-relevant-sounding appeal to authority whose persuasive force (it sounds authoritative and well-formed) far exceeds its logical force
This sourceDocuments large-language-model reasoning that is fluent and plausible yet unfaithful/content-defective, generating well-formed arguments whose stated grounds do not actually warrant the conclusion — the form-passes-content-fails structure in generated reasoning.
- In an AI deployment, a language model is asked whether a policy is sound and returns a fluent, confident argument that turns on the claim "every credible economist agrees," a fabricated-but-relevant-sounding appeal to authority whose persuasive force (it sounds authoritative and well-formed) far exceeds its logical force
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Links previously used in the corpus¶
Before the registry existed this work was also linked 1 other way.
Registry ID ref:9bbfe4fa685a · see in the full table