Rendered at 12:45:34 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Planktonne 2 days ago [-]
Of course not. Because the article uses the words 'thought' and 'reasoning' and even 'faithful' to mean something other than their normal meanings, but then expects them to behave exactly the same.
Every field has terms of art, and 'reasoning' is one for LLMs. But that doesn't mean it has the same properties as 'reasoning' in other contexts, because you're not referring to the same thing.
Why doesn't my asteroid belt buckle?
stymaar 2 days ago [-]
Related: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces![1]
> Our findings consistently challenge the prevailing narrative that intermediate tokens constitute a semantically meaningful reasoning process. First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks.
The author of this paper is in ASU and does a lot of excellent work in this space. People should check it out. Especially the paper titled "Beyond Semantics..." His twitter is also active _and_ high SNR.
Over time, I've learned to accept that many people -- even very clever ones -- are incapable of holding a metaphor at arm's length. Once they accept the words of a metaphor as applicable at all, the metaphor collapses entirely into literalism for them. They can no longer see that the metaphor was just a tool with inherent limitatation and boundaries.
Because the field of artificial "intelligence" is constructed around the idea of applying psychological metaphors to computational systems (a very powerful idea!) it's almost a worst case scenario for these people.
Suddenly, they're reversing the metaphors and applying computational schema to psychological processes ("aren't we really just stochastic parrots ourselves?!"); or, like here, they find themselves surprised and confused when they stumble across the natural boundaries of the metaphor experimentally.
It's because they never had sight of the boundaries in the first place and maybe never can quite see them. The words only make sense to them as literal equivalence, and so their surprise when they run into stuff like this is earnest and deep.
empath75 2 days ago [-]
I don't think the boundary between "generalization" and "metaphor" is very well defined. When you go from an exemplar of 1 to 2, you're going to find all kinds of edge cases where attributes of the thing being demonstrated that had seemed to be essential turn out to not be necessary.
I think you certainly could look at LLMs as "thinking" metaphorically, but I also don't think it is necessarily only a metaphor.
swatcoder 2 days ago [-]
Well, it's unusual to "generalize" an idea if you only had two exemplars, one so old and so complicated that all your terms are specifically referent to it and often even hard to be precise about; and the other is both extremely novel and plainly distinct in both its mechanisms and behaviors.
While maybe that boundary can be fuzzy, we're unequivocally and deeply in "metaphor" territory here.
florianherrengt 2 days ago [-]
This paper puts words to something I’ve noticed repeatedly with LLMs, particularly Qwen3.6. When I read its reasoning, it appears to recognise the mistake and then carry on as if it hadn’t noticed it at all.
> models often determine their answers based on implicit biases tied to question templates, then construct reasoning chains to justify their predetermined conclusions
> its reasoning was correct right until the final step (Yes/No answer)
flyingpumba 2 days ago [-]
First author here, surprised to see the paper in HN! :)
When doing the paper we noticed that models are very good at generating post-hoc plausible CoT, which to me knowledge can happen quite often with relatively easy tasks.
Must we always see this restated every time? It's getting a bit stale always seeing these kinds of comments on articles about LLM.
ethin 2 days ago [-]
I agree, and I very strongly dislike it, to be polite about it. It contributes absolutely nothing and is an excellent way of hand-waving away literally anything an AI model does. Saying "well people do this too" is a great way to rationalize away anything you can imagine that an AI model would be capable of, because "humans do it too so what's the big deal, guys?"
uludag 2 days ago [-]
I can just immagine the response to a headline "LLM chooses mass death: thousands killed in horrific AI accident" being something like "lots of humans have caused mass death too."
Kim_Bruning 23 hours ago [-]
Maybe we can think up a (personal) rule we can follow?
A naked "natural intelligences do this too" might be a bit too short to be useful. But if we can add when/where, cite papers, or show ways in which the parallel operates, then it might be useful.
Compare, eg, talking about a robot arm, and someone goes "a natural arm does this too". You can tell about the fact that it has the same degrees of freedom in the same places, or how this pertains to inverse kinematics, or etc...
Same way here, "this happens to be how natural intelligences seem to solve this too! According to Foo, Bar, Baz et al (2026) the gadget is always twiddled beforehand in macaque apes. " or "Same for natural intelligence: I've noticed I use the same general algorithm myself. I've always considered this the correct way to do translation between languages".
--
A more concrete example of a useful answer here.
Natural intelligence does this too! When given the question "explain your reasoning" humans are indeed quite prone to post-hoc confabulation. [1]
It's the grounded portion of a feedback loop searching for the 'why is this happening' thinking. I'd imagine most of my own comments in this area boil down to "GIGO" most of the time.
It's a relevant comment in this instance because we're discussing concepts you need to be both trained and practiced in to reason about, and that our discipline has traditionally been blind to. Plenty of people working with LLM context issues who've never been exposed to the idea of 'subtext' or could tell you why it would matter to their direction of effort.
8note 2 days ago [-]
continuing with the "we need open training data" thread
how much of the training data had thinking traces that dont make sense to people as being actually a description of why the output should be that way?
8note 2 days ago [-]
or rather, its not productive to "we should do better with artificial intelligence"
cyanydeez 2 days ago [-]
You think, "this problem" is qn LLM problem?
2 days ago [-]
phailhaus 2 days ago [-]
No they don't, human intelligence has the ability to form an internal model of itself, which allows it to "notice" its own mistakes and change.
cyanydeez 2 days ago [-]
Many who watched the last decade knows just because its possible to noticed mistakes and change, its clearly not a reliable process.
elictronic 2 days ago [-]
The current political climate is well reasoned and intentional. It might not be yours or mine, however the system is working exactly as the ones paying for it have intended.
delichon 2 days ago [-]
My brother and I have been arguing about that all of our lives. He believes everything is intentional and it's just a matter of discovering who benefits. I see chaos that nobody intends or controls. His political landscape is a tapestry of conspiracy theories and mine is a fog of war. I think his is more comforting, since it admits a possibility of a rational, predictable world.
Terr_ 2 days ago [-]
I'd synthesize those as: "There is a lot of chaos with no central plan, but every small piece happens because someone believes they will benefit."
In other words, a lot of this depends on what scale/scope is being inspected. On the high level, the world is chaos rather than a meticulous and inscrutable plan of the Illuinati Shadow Cabal. On the low level, people do things for reasons, even if they're dumb ones.
With respect to the "current political climate", I'd like to suggest that a lot of dumb or seemingly "against their own interests" stuff is due to people prioritizing costly in-group loyalty signals. Their interest in staying good with the tribe is just higher than their interest against a dumb national policy.
dgellow 2 days ago [-]
It’s intentional in the sense that actors are acting intentionally for their own benefit (or at least what they believe is beneficial) and following incentives. Not that there is a master planner who manipulates everything
elictronic 2 days ago [-]
The current admin seems to have quite a few long term plans they have been working towards.
Project 2025, Maralago accords.
So far the only major policy item the Trump admin seems to have not intended was the Iran War. Israel killing the intended replacement, Iran leveraging the straight of Hormuz, and dropping three Tomahawks on an elementary school really botched that one.
dgellow 2 days ago [-]
Yeah, if we are talking about the Trump admin it’s definitely a conspiracy. Pretty much Peter Thiel’s cabal. But they are pretty open about their plans
cyanydeez 2 days ago [-]
the thing is, that fog might've been true decades ago, but for 100 billionaires to sit in a virtual smokey room and do the shit they want to do, that's not really a conspiracy.
It's just peter theil's texting groups.
The ability to conspiracy both willing and unwilling is such a low threshold now, it's virtually indistinguishable.
You watch one billionaire do something and you're like, I'm a billionaire, I should do that too.
The fact that there's so few billionaires, the probability that they conspire together both direct and indirect approaches 1.
The inverse of course is rediciously hard to conceive: the working class bands together to get something like universal healthcare.
freejazz 2 days ago [-]
Yeah and it's not great then either
ethin 2 days ago [-]
[flagged]
paimapi 2 days ago [-]
[dead]
bee_rider 2 days ago [-]
I see this when just using some chat bot that shows the “reasoning” steps (ad-hoc observation of course, it’s really cool that people are actually studying it).
It is annoying when the bot seems be “reasoning” correctly and then makes an obvious mistake at the end. And perplexing when it seems to be completely wrong and then pull the right answer out of a magic hat at the end.
I guess it makes sense; the “reasoning” steps aren’t actually doing logic, just adding more context to influence the final generation, right? But it is weird to see.
sergio_valencia 2 days ago [-]
This paper made me wonder not whether the chain we can read is actual “thought,” but what conclusions we can draw by observing it. It is an output channel, but not direct access to the black box that creates it. The question pairs present an interesting experiment, but they also got me thinking about semantics: the same underlying relation can have many valid representations, and how models reason across those representations can tell us more about their stability and correctness. Basically, are models semantically consistent when given different representations of the same underlying relation? Do their answers transform as the relationship requires, and do their explanations remain consistent with that relationship?
There's no reason to believe the model's self-reported "thinking" bears any relation to the mechanics by which it arrived at some output.
orbital-decay 2 days ago [-]
It's... complicated. Yes, RL reward hacking makes it learn "bird language" and yes, reasoning traces can be misleading. However they also pretty clearly steer the final reply and not simply justify it, and can stay somewhat coherent and relevant with readability SFT and rewards. All these phenomenas coexist, they aren't mutually exclusive. Reasoning traces are still useful for debugging.
8note 2 days ago [-]
that sounds testable - if you skip the reasoning tokens, do you get the same result?
if not, then there's certainly some bearing, but not necessarily in how we read the tokens as text
rcxdude 2 days ago [-]
It is and has been - skipping the tokens causes performance to drop. But replacing the tokens with filler causes the performance to drop, but by a lot less. You can also train models to emit broken or unrelated thinking tokens and their performance is also not much worse than the ones that are trained to output somewhat coherent thinking traces.
This points to a hypothesis that the content of the tokens is only slightly related to the mechanism by which it improves performance, and that primarily the extra tokens allow the original prompt to be processed more deeply by the model, because earlier tokens will essentially pass through the model many more times than later ones.
neuroticnews25 1 days ago [-]
I'll never understand some downvoters here.
ForHackernews 1 days ago [-]
AI boosters denying that their machine god is burbling sweet nothings to itself?
2 days ago [-]
kibwen 2 days ago [-]
"Study: Communing With The Gods of Mount Olympus Via the Oracle at Delphi Is Not Always Faithful"
Every field has terms of art, and 'reasoning' is one for LLMs. But that doesn't mean it has the same properties as 'reasoning' in other contexts, because you're not referring to the same thing.
Why doesn't my asteroid belt buckle?
> Our findings consistently challenge the prevailing narrative that intermediate tokens constitute a semantically meaningful reasoning process. First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks.
[1]: https://arxiv.org/abs/2504.09762
Related: Poster side dialogue and Q&A about this work at ICML. Very good. https://news.ycombinator.com/item?id=49277303
Because the field of artificial "intelligence" is constructed around the idea of applying psychological metaphors to computational systems (a very powerful idea!) it's almost a worst case scenario for these people.
Suddenly, they're reversing the metaphors and applying computational schema to psychological processes ("aren't we really just stochastic parrots ourselves?!"); or, like here, they find themselves surprised and confused when they stumble across the natural boundaries of the metaphor experimentally.
It's because they never had sight of the boundaries in the first place and maybe never can quite see them. The words only make sense to them as literal equivalence, and so their surprise when they run into stuff like this is earnest and deep.
I think you certainly could look at LLMs as "thinking" metaphorically, but I also don't think it is necessarily only a metaphor.
While maybe that boundary can be fuzzy, we're unequivocally and deeply in "metaphor" territory here.
> models often determine their answers based on implicit biases tied to question templates, then construct reasoning chains to justify their predetermined conclusions > its reasoning was correct right until the final step (Yes/No answer)
When doing the paper we noticed that models are very good at generating post-hoc plausible CoT, which to me knowledge can happen quite often with relatively easy tasks.
You might be interested in reading this other paper that came out after ours: https://arxiv.org/abs/2507.05246
A naked "natural intelligences do this too" might be a bit too short to be useful. But if we can add when/where, cite papers, or show ways in which the parallel operates, then it might be useful.
Compare, eg, talking about a robot arm, and someone goes "a natural arm does this too". You can tell about the fact that it has the same degrees of freedom in the same places, or how this pertains to inverse kinematics, or etc...
Same way here, "this happens to be how natural intelligences seem to solve this too! According to Foo, Bar, Baz et al (2026) the gadget is always twiddled beforehand in macaque apes. " or "Same for natural intelligence: I've noticed I use the same general algorithm myself. I've always considered this the correct way to do translation between languages". --
A more concrete example of a useful answer here.
Natural intelligence does this too! When given the question "explain your reasoning" humans are indeed quite prone to post-hoc confabulation. [1]
[1] https://home.csulb.edu/~cwallis/382/readings/482/nisbett%20s... "Telling More Than We Can Know: Verbal Reports on Mental Processes" (this citation is quite old and may have been superseded, mostly just to illustrate how the rule might work)
It's a relevant comment in this instance because we're discussing concepts you need to be both trained and practiced in to reason about, and that our discipline has traditionally been blind to. Plenty of people working with LLM context issues who've never been exposed to the idea of 'subtext' or could tell you why it would matter to their direction of effort.
how much of the training data had thinking traces that dont make sense to people as being actually a description of why the output should be that way?
In other words, a lot of this depends on what scale/scope is being inspected. On the high level, the world is chaos rather than a meticulous and inscrutable plan of the Illuinati Shadow Cabal. On the low level, people do things for reasons, even if they're dumb ones.
With respect to the "current political climate", I'd like to suggest that a lot of dumb or seemingly "against their own interests" stuff is due to people prioritizing costly in-group loyalty signals. Their interest in staying good with the tribe is just higher than their interest against a dumb national policy.
Project 2025, Maralago accords.
So far the only major policy item the Trump admin seems to have not intended was the Iran War. Israel killing the intended replacement, Iran leveraging the straight of Hormuz, and dropping three Tomahawks on an elementary school really botched that one.
It's just peter theil's texting groups.
The ability to conspiracy both willing and unwilling is such a low threshold now, it's virtually indistinguishable.
You watch one billionaire do something and you're like, I'm a billionaire, I should do that too.
The fact that there's so few billionaires, the probability that they conspire together both direct and indirect approaches 1.
The inverse of course is rediciously hard to conceive: the working class bands together to get something like universal healthcare.
It is annoying when the bot seems be “reasoning” correctly and then makes an obvious mistake at the end. And perplexing when it seems to be completely wrong and then pull the right answer out of a magic hat at the end.
I guess it makes sense; the “reasoning” steps aren’t actually doing logic, just adding more context to influence the final generation, right? But it is weird to see.
From March last year: https://transformer-circuits.pub/2025/attribution-graphs/bio...
There's no reason to believe the model's self-reported "thinking" bears any relation to the mechanics by which it arrived at some output.
if not, then there's certainly some bearing, but not necessarily in how we read the tokens as text
This points to a hypothesis that the content of the tokens is only slightly related to the mechanism by which it improves performance, and that primarily the extra tokens allow the original prompt to be processed more deeply by the model, because earlier tokens will essentially pass through the model many more times than later ones.