Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm using 'reason' in the language-captured engineering sense. They generate text as-if they reason. Those reasoning traces poorly correlate with their given answers (which is called "cheating" by people who fall for the illusion). It is in part because the reasoning doesn't entail their answers but merely 'steers the text as-if it does' that is fatal for calling it reasoning. The process by which reasoning traces and user-facing completions are generated isnt reasoning. Reasoning is a specific process, and LLMs don't do it.

Now, of course, humans can also generate answers without reasoning too -- and in those cases, that isnt reasoning also. And in cases where people confabulate, that isnt reasoning likewise. But humans, and many classes of animals, do reason. They do reach answers via inferential entailments, not merely steered correlations.

LLMs provide imitations of arbitrary mental capacities "in the text domain", ie., the generate text as-if the LLM had those capacities. Insofar as the text generated is useful, for an engineer, that's sufficient.

As a person with scientific commitments to reality rather than its immitation, i retain the ordinary non-engineered meanings of these terms: reasoning is a deterministic inferential process over propositions; and a reasoning agent is one which has the capacity to represent propostions and their entailments, and does so when they reason. LLMs fail at all hurdles here: they have no propositonal states (ie., no rich representations), no inferential process which unites them, and so on.

You can always get abitarily close to appearing as-if, if the LLM is trained on a vast number of reasoning examples, of course. But as I said, you still have the "stochastic parrot" problem. Now your problem is your reasoning is parroted. This is a nice problem to have, if you're just playing chess -- but is a catastrophic problem if you're hacking civil infrastructure.

 help



But how do you falsify that?

If LLM can always imitate closely enough to appear as-if, how can you ever separate it from whatever actual intelligence is?


Well intelligence is not measured by patterns in text. The illusion only takes place in the text domain.

Even then, it's a pretty fragile illusion at the moment. Clearly the reasoning traces dont ground the answers. There's no intelligence taking place even as-measured by text.


Why could you not use text to measure intelligence? Isn't that what we used to do? Have students write an essay, and then deduce that they had enough intelligence to put some thoughts together?

Well that's the trick. That the systems we use, as a proxy, to measure intellinnce in people are actually fairly easy to immitate.

But let's be clear these were always, and are, bad measures of intelligence. You cannot test a dolphin this way. And its easy to cheat on tests either thru recall , wrote-learning, etc. and IQ tests haev very poor individual test-retest reliability.

In humans there's a convenient correlation that verbal articulation in text is a strong but weak correlate of intelligence. Its "Good enough" for allocating meat bodies to our various institutions. But if you've met many well-tested people you'll realise how, in practice, terrible this measuring approach is. The world we inhabit is filled with misclassified "intellects" who perform well under text-based rubrics. Add LLMs to that heap, the cheater par execellence.


How can you know that verbal articulation is a weak correlate? There must be some other, canonical way to measure intelligence?

Sure. Let's first establish a measure is not what is being measured. Then, as far as measures go, we need measures that capture the entire animal kingdom. I can go it into it, but its a lot of time and I'm busy atm.

Consider thought that all mammals have imagination, and model-based reinformcement, and a wide vareity of other capacities required for intelligence. And so merely issuing "text" captures, incorrelate, only these capacities by proxy.

I'm sure if you thought about it yhou could come up with tests that distinguish lizards from birds and the greater apes from the lesser. Those are the tests




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: