> They are also notably bad at judging the historical significance of what they find.
I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.
It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
> seven chord groups
This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.
AI capabilities are "spiky": they extend far in some dimensions but fall short in others, seemingly at random. See for example the recent "thus spoke compute" musical[1]. It's an absolute banger, the graphics are impressive, and so is the writing. But some of the metaphors make no sense, the text highlights are in the wrong places, and the train animation at 2:35 is running backwards!
A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.
No, one might not say that. Calculators are not regarded as even a Narrow Intelligence because there's no intelligence. And no, not because of the 'humans so special' or 'it's software!' tautology that oft gets repeated in these discussions. I mean there's no adaptability whatsoever. A Chess bot has it (in its narrow domain of Chess). A calculator does not.
My point (using irony) is that "spiky intelligence" is a somewhat hollow phrase, something that sounds like it could be objective but really it's just a flavor on top of "I'll know it when I see it."
Take anything "intelligent", alter it to be spikier and spikier, and eventually *poof* somehow the intelligence vanishes. You can do the same with the phrases "flawed intelligence" or "specialized intelligence."
> Calculators [have] no intelligence. [...] I mean there's no adaptability
To short-circuit a long discussion, I submit that "adaptability" will (once the Scooby Doo gang catches it) turn out to be "intelligence" in a tautological mask, both equally undefinable except in relation to one-another.
Something will be intelligent because you perceive adaptability, and it'll be adaptable because you infer intelligence. If it doesn't seem adaptable, it can't be intelligent, and if you don't want it to be intelligent, it won't have "real" adaptability.
> A Chess bot has [adaptability] (in its narrow domain of Chess). A calculator does not.
My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?
>My (ironic) point is that "spiky intelligence" seems rather unhelpful because the "intelligence" part still operates under the rules of "I just know it when I see it."
All intelligence is 'spiky' or 'jagged' or whatever, even human intelligence. We evolved in certain environments and situations and sometimes there's a mismatch and it causes all sorts of wonky things. We just call them funny names like optical illusions and cognitive biases. But it's the same thing. I agree there's no point in call LLMs a 'spiky' intelligence, but for probably the oppoosite reasons as you.
>My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?
A chess engine can be dropped into a board position it has never encounterd and search over possible continuations, evaluating and selecting actions based on the state it finds itself in.
A calculator solving x+3=7 is doing something quite different. The fact that x can take arbtrary values doesn't make the calculator adaptive; it just means the fixed procedure operates over a range of inputs. Every problem a calculator can solve was effectively anticipated when it was built, and anything outside that grammar produces an error with no partial credit. A chess engine can face positions nobody enumerated, in situations nobody could even dream of and still produce sensible moves.
The core of being intelligent is being able to make decisions independently. That's why we hire smart people - to make better decisions. You can't make decisions if everything is spelt out for you. And naturally if you can't make decisions you can't adapt.
This analogy may be too close to the real thing to work, but it reminds me of a Chinese room type situation where its entire understanding of the world is through messages of text.
You say that’s an error a human couldn’t do, but imagine if the human has never seen or touched the kind of item you were making and relied entirely on text descriptions to build its ontology. Off by 90 seems like such a believable mistake.
There's a famous early world map created by Ptolemy from compiling reports of sailors which is amazingly accurate for its time but had the country of Scotland off by 90 degrees.
Also, not to sound like a naive hypemonger, but: in a decade I'd bet a ton of money the best AI systems will make strange mistakes of this nature at a far, far lower rate than they do today. They will gain a more holistic and more human-like perspective about each task.
(even if it's through some silly means like explicitly talking to themselves like "if I were a human doing this, what [... 5 million tokens in 2 seconds ...]" but also of course if they crack ASI and get something more efficient and intelligent than a human brain by then)
I think of it kind of like how Chess AI make "mistakes" which are unrecognizable to humans but a stronger AI would be able to pick them apart. That's kind of scary...
I agree, there’s the collective intelligence that created all the content used to train the model. The model is a superposition of all that material with RL tuning. Analogously to reading a book, the intelligence you perceive is from the book’s creator.
What about DNA? There are things we do that we never read in a book, maybe never seen someone else do them vefore, but we still do them. Or we still feel a certain way. That doesn't come from "human training data", unless you count the DNA as training data.
In general llms are weak with spatial reasoning. This seems to be an unsolved problem. Probably because human language is generally imprecise spatially and humans think about spatial problems in visual terms. I wonder if having an llm make a 3d design in a format an image model could check would result in a better outcome?
Having used them heavily for 3D model understanding since February I can say it's nuanced.
Opus 4.6 and 4.7 were bad, but GPT 5.2 and above were very usable. Opus 4.8 was usable, but the GPT 5.x series was better.
Fable is great.
Opus 5.0 was interesting. It could solve some problems that Sol 5.x couldn't solve (applying a G2 curve on a 3 way corner where one face was a Bezier curve) but you had to be super prescriptive ("only answer the question"/"only do what I tell you and stop when done") or it would go on a hugely involved validation journey that didn't really achieve a lot.
Opus 5.5 is better than that was in that respect.
Sol 6.x is great, and my daily driver for this (I use Opus for coding though)
Astra can solve problems that Sol can't but for some reason on easy stuff makes uglier solutions.
For all models it's very interactive though - we aren't at the "agentic design" phase for most things yet.
I think that this is an unsolved problem in the same way that mangled fingers in image generation was an unsolved problem.
Through at least Opus 4, LLMs were practically useless for authoring any sort of coherent procedural closed-curve geometry (I know this with strong confidence because of the little animated guys at https://letterspractice.com).
Opus 5.5 can bang it all out. Possibly a deliberate RL sort of thing or maybe another surprise emergent capability.
Is it that LLMs are weak with spatial reasoning (and memory) or is it that we are unusually good at it?
When I need to use a program I seldomly use I'm far more likely to remember where I need to click to open it than the word I need to search for to open it.
Yes I like to think of humans with built-in accelerators for certain tasks -- our visual and spatial reasoning is off the charts presumably because it's a life or death skill!
Good read! How many pages of text were in scope? I'm not sure if the 1615+1629 pages were the total or just a subagent.
If they were the total I would say it was arguably more impressive the author was able to narrow it down to just 3000 pages than it was to find the dodo mention amongst those!
My dad died and had many many notebooks of his journals with very hard-to-read handwriting. Is it worth the effort to scan all of these so I can feed them in and go to work. Seems like so much minutia is out there, ready to be meta-understood.
I set up an OCR flow using local models on all my many tens of journals stretching back the last 30 years.
I would say it's about 80% accurate, which means it's missing enough key words to make a lot of it uselessly unintelligible. I can easily compare the images against text I turn up in a grep which is nice if I'm looking for something.
Allegedly Claude set up a system for retraining for my handwriting, but it would require me to manually revise several hundred pages by hand so I don't think I'll ever do it.
> Epistemological weirdness
> They are also notably bad at judging the historical significance of what they find.
I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.
It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
> seven chord groups
This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.
AI capabilities are "spiky": they extend far in some dimensions but fall short in others, seemingly at random. See for example the recent "thus spoke compute" musical[1]. It's an absolute banger, the graphics are impressive, and so is the writing. But some of the metaphors make no sense, the text highlights are in the wrong places, and the train animation at 2:35 is running backwards!
A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.
[1] https://www.youtube.com/watch?v=Cq8qO-NjYIg
One might say a calculator is just another example of "spiky intelligence", merely spikier.
No, one might not say that. Calculators are not regarded as even a Narrow Intelligence because there's no intelligence. And no, not because of the 'humans so special' or 'it's software!' tautology that oft gets repeated in these discussions. I mean there's no adaptability whatsoever. A Chess bot has it (in its narrow domain of Chess). A calculator does not.
My point (using irony) is that "spiky intelligence" is a somewhat hollow phrase, something that sounds like it could be objective but really it's just a flavor on top of "I'll know it when I see it."
Take anything "intelligent", alter it to be spikier and spikier, and eventually *poof* somehow the intelligence vanishes. You can do the same with the phrases "flawed intelligence" or "specialized intelligence."
> Calculators [have] no intelligence. [...] I mean there's no adaptability
To short-circuit a long discussion, I submit that "adaptability" will (once the Scooby Doo gang catches it) turn out to be "intelligence" in a tautological mask, both equally undefinable except in relation to one-another.
Something will be intelligent because you perceive adaptability, and it'll be adaptable because you infer intelligence. If it doesn't seem adaptable, it can't be intelligent, and if you don't want it to be intelligent, it won't have "real" adaptability.
> A Chess bot has [adaptability] (in its narrow domain of Chess). A calculator does not.
My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?
>My (ironic) point is that "spiky intelligence" seems rather unhelpful because the "intelligence" part still operates under the rules of "I just know it when I see it."
All intelligence is 'spiky' or 'jagged' or whatever, even human intelligence. We evolved in certain environments and situations and sometimes there's a mismatch and it causes all sorts of wonky things. We just call them funny names like optical illusions and cognitive biases. But it's the same thing. I agree there's no point in call LLMs a 'spiky' intelligence, but for probably the oppoosite reasons as you.
>My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?
A chess engine can be dropped into a board position it has never encounterd and search over possible continuations, evaluating and selecting actions based on the state it finds itself in.
A calculator solving x+3=7 is doing something quite different. The fact that x can take arbtrary values doesn't make the calculator adaptive; it just means the fixed procedure operates over a range of inputs. Every problem a calculator can solve was effectively anticipated when it was built, and anything outside that grammar produces an error with no partial credit. A chess engine can face positions nobody enumerated, in situations nobody could even dream of and still produce sensible moves.
The core of being intelligent is being able to make decisions independently. That's why we hire smart people - to make better decisions. You can't make decisions if everything is spelt out for you. And naturally if you can't make decisions you can't adapt.
Thermostats are smarter than calculators :thinking_face:
They adapt to the buttons you press. Thus, intelligent and capable of feeling pain.
How often does the calculator get something wrong?
This analogy may be too close to the real thing to work, but it reminds me of a Chinese room type situation where its entire understanding of the world is through messages of text.
You say that’s an error a human couldn’t do, but imagine if the human has never seen or touched the kind of item you were making and relied entirely on text descriptions to build its ontology. Off by 90 seems like such a believable mistake.
There's a famous early world map created by Ptolemy from compiling reports of sailors which is amazingly accurate for its time but had the country of Scotland off by 90 degrees.
Also, not to sound like a naive hypemonger, but: in a decade I'd bet a ton of money the best AI systems will make strange mistakes of this nature at a far, far lower rate than they do today. They will gain a more holistic and more human-like perspective about each task.
(even if it's through some silly means like explicitly talking to themselves like "if I were a human doing this, what [... 5 million tokens in 2 seconds ...]" but also of course if they crack ASI and get something more efficient and intelligent than a human brain by then)
I think of it kind of like how Chess AI make "mistakes" which are unrecognizable to humans but a stronger AI would be able to pick them apart. That's kind of scary...
Unfortunately, there are reports that they have “dumbed” down Opus 5.5 already.
https://github.com/ninjahawk/livenerf
Which reports? There are lots of people watching model quality now, so it seems like there would be clear evidence if it happened already.
It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
Maybe the model isn't intelligence in any form, except perhaps as an imperfect reflection of the intelligence of its training data.
I agree, there’s the collective intelligence that created all the content used to train the model. The model is a superposition of all that material with RL tuning. Analogously to reading a book, the intelligence you perceive is from the book’s creator.
What about DNA? There are things we do that we never read in a book, maybe never seen someone else do them vefore, but we still do them. Or we still feel a certain way. That doesn't come from "human training data", unless you count the DNA as training data.
Palaeolithic natural selection did the training.
In general llms are weak with spatial reasoning. This seems to be an unsolved problem. Probably because human language is generally imprecise spatially and humans think about spatial problems in visual terms. I wonder if having an llm make a 3d design in a format an image model could check would result in a better outcome?
Having used them heavily for 3D model understanding since February I can say it's nuanced.
Opus 4.6 and 4.7 were bad, but GPT 5.2 and above were very usable. Opus 4.8 was usable, but the GPT 5.x series was better.
Fable is great.
Opus 5.0 was interesting. It could solve some problems that Sol 5.x couldn't solve (applying a G2 curve on a 3 way corner where one face was a Bezier curve) but you had to be super prescriptive ("only answer the question"/"only do what I tell you and stop when done") or it would go on a hugely involved validation journey that didn't really achieve a lot.
Opus 5.5 is better than that was in that respect.
Sol 6.x is great, and my daily driver for this (I use Opus for coding though)
Astra can solve problems that Sol can't but for some reason on easy stuff makes uglier solutions.
For all models it's very interactive though - we aren't at the "agentic design" phase for most things yet.
Here's a sample of what I've been able to get them to design with me: https://x.com/nlothian/status/2099023496794018067
I think that this is an unsolved problem in the same way that mangled fingers in image generation was an unsolved problem.
Through at least Opus 4, LLMs were practically useless for authoring any sort of coherent procedural closed-curve geometry (I know this with strong confidence because of the little animated guys at https://letterspractice.com).
Opus 5.5 can bang it all out. Possibly a deliberate RL sort of thing or maybe another surprise emergent capability.
Is it that LLMs are weak with spatial reasoning (and memory) or is it that we are unusually good at it?
When I need to use a program I seldomly use I'm far more likely to remember where I need to click to open it than the word I need to search for to open it.
Yes I like to think of humans with built-in accelerators for certain tasks -- our visual and spatial reasoning is off the charts presumably because it's a life or death skill!
Good read! How many pages of text were in scope? I'm not sure if the 1615+1629 pages were the total or just a subagent.
If they were the total I would say it was arguably more impressive the author was able to narrow it down to just 3000 pages than it was to find the dodo mention amongst those!
Those are years, not page counts, right? The article mentions "millions of records", but I'm not sure how big a record can be.
My dad died and had many many notebooks of his journals with very hard-to-read handwriting. Is it worth the effort to scan all of these so I can feed them in and go to work. Seems like so much minutia is out there, ready to be meta-understood.
I set up an OCR flow using local models on all my many tens of journals stretching back the last 30 years.
I would say it's about 80% accurate, which means it's missing enough key words to make a lot of it uselessly unintelligible. I can easily compare the images against text I turn up in a grep which is nice if I'm looking for something.
Allegedly Claude set up a system for retraining for my handwriting, but it would require me to manually revise several hundred pages by hand so I don't think I'll ever do it.
https://github.com/508-dev/journal-ocr
What a great read, I generally associate substack with verbose, low quality content, but definitely not the case here!