I've used Astra, Fable 5.1, Sol 5.6, Opus 5. They're definitely making progress and the proof is in the pudding - I reach for the most advanced model as much as I can. But I wouldn't say their capabilities have increased dramatically. I use it for both coding and non-coding workloads and I don't feel I can do anything dramatically different. Again at the end of the day this is a debate on semantics because we can have very different definitions of "dramatic improvement".
(I'm not factoring the benchmarks into the discussion, because I've never quite cared about them)
I really enjoy the balance of speed and accuracy of Astra. I can definitely see it become my driving model for most tasks, technical and non-technical.
However, I don't see it as such a massive leap compared to Fable or Sol. As ever, there's a mismatch between the benchmarks and my daily experience of the models.
What do you all think about Astra now that it's been out for a few weeks?
> What do you all think about Astra now that it's been out for a few weeks?
Best model put out so far by any of the frontier labs. Way better than Anthropics models, especially in actual text generation. Claudes fodder heavy text is ridiculous.
> However, I don't see it as such a massive leap compared to Fable or Sol.
It's hard to quantify these things without burning tons of tokens. But Fable has been a huge disappointment for me with the sole exception of graphics (UI/GPU shaders). It burns an obscene amount of tokens and barely produces output better than Opus 5.
Edit because I forgot to mention that Fable is the only modern model that seems to splat out random Chinese or Arabic glyphs. And 5.1 does it more than 5
Such a massive leap at averaging possible use cases from previous data collected.
My guess is : collect all the prompt and their satisfaction score. group them by similarity . For each group pretrain the next model on that . Get these results ready.
Next model generation feed them back those answers.
I still develop in smaller chunks, checking nearly all the output. However I have a work project (building the warehouse and BI for a client) that is well-specified and where I will try to few-shot the development. Hope it delivers.
How do you normally verify the work product of a "few-shot" development process? Do you scrutinise the source code like with human developers? Or do you just run the test suite and click around the app to check if it seems to work?
I haven't done it in a production project - this will be the first time for me. I have specified the architecture and data definitions pretty well. The tests will be run against the customer's Excels, which is what the warehouse will be replacing. I'll check the general shape of pipelines, models, orchestration code, etc. but in many parts I probably won't review the code myself.
Like I said in another comment, this famous article by the Atlantic [1] has pretty convincing testimony that kids studying the humanities at elite universities can't read either.
I agree with the sentiment that we tend to misrepresent the culture of mass societies, as if at one point in time there were droves of people reading Aristotle. But the ability to read and our transition to an oral or post-literate culture is very much affecting elites too. There is pretty good evidence when you compare contemporary histories. Or if you live in a bourgeois milieu, you just need to ask around.
I don't think that reading necessarily makes you a better person. I've met a good amount of people of the highest moral dignity who do not read. But on the whole, reading helps build competence - moral and technical - and create a sense of shared culture. Going back to an "oral culture" is nothing but a regression.
One of the reasons why we read less long-form, according to a bunch of cognitive science studies and to my own anecdotal experience, is because our attention is shot.
A viral article in The Atlantic [1] shows evidence and testimony that kids at elite universities simply can't read anymore. I've seen the claim repeated in many other places and looking around me I believe it. Those kids are not replacing Dostoevsky or Aristotle with more compact versions of it.
The airport book industry is indeed full of books that could be fit in ten pages. But there's a universe of useful - potentially life-changing - books beyond that, and not only in the humanities. Some modern technical books justify their length. I seriously doubt that something like The Data Warehouse Toolkit can be replaced by reading piecemeal blogposts online.
Same goes for history, I sincerely doubt you can properly learn it in depth without reading long-form. In my experience, people who only listen to history podcasts or consume 10-minute videos tend to have a much poorer command of the facts and of the broader motions of history.
>Many of [students] do very well—but a striking share perform abysmally (see chart 1). Across rich countries some 8% of students in tertiary education notch up a score in literacy no better than one might expect from a ten-year-old child. The share is about the same for numeracy. Worse, the share at or below this bar has risen since the tests were last run, a little over a decade ago. The share of very poor performers in literacy has more than doubled.
Is this your idea of an intelligent contribution to the conversation?
There is plenty of evidence in the article from real professors at real universities (ex: Princeton), saying that they’ve had to dramatically cut reading workloads from humanities courses. I’ve heard similar things from European universities.
Other hard evidence is in plummeting reading comprehension scores in many countries that participate in PISA.
In the States, there is plenty of other stats that point that students are below their expected reading grade - the worse in generations.
And all you have to contribute is that this is “idiotic”? The irony. Inform yourself.
Above all, NONE of that excuses the claim that "that kids at elite universities simply can't read anymore". That one is simply, factually not true. There is such a thing as not being able to read. Students at
elite universities, who btw are not kids, dont fit there.
Second, yes, and mostly because I did actually read those stats and
details of how they are made. All the nuance from the actual
scientific article gets exaggerated into completely different claims
just to create outrage and perpetuate it again and again. Usually by
people who do not know what PISA or reading grades measure.
Reading comprehension as measured by PISA have zero to do with long
form reading of books including classics. It is not trained by reading
those nor measure reading those. The short for articles full of
graphs, scientific claims and such the Atlantic complains about is how
you train people for PISA. I am not criticizing PISA here btw. It is
ok to measure that.
> In the States, there is plenty of other stats that point that students are below their expected reading grade - the worse in generations.
Interestingly, those who made this claim usually do not know how
reading grades are measured or even how high they go. "The worse in
generations" claim is especially dubious and tend to based on
comparison between "guess performance of imaginary past elite student
in undefined historical period" with low performing student today. Low
performing students 3-4 generations ago disappear, because we imagine
them to not exist.
There is outrage every time a school replaces one book by another. Or
when a course that has no reason to assign full Odyssey assigns range
of materials with one chapter from it. There is outrage when student
names a contemporary book as the favorite one instead of a classic.
There is more to literacy than being able to literally say the words that are represented by the letters on the page.
Your argument here, funnily enough, is an example of the exact kind of literacy decline as stated in the OP article.
> Allusions are gone; metaphor is treated literally, often provoking dim-witted outrage.
No, professors are not claiming that college students are literally incapable of sounding out words.
They are saying that students are not able to keep attention on large bodies of text, to follow arguments, understand authorial intent, describe themes, identify rhetorical techniques, or intuit metaphors (such as this very use of reading as synecdoche/metaphor for this whole class of literacy skills which is so obvious that it feels absurd to even have to explain it)
An 8-year-old may be able to read as proven by their ability to answer basic plot questions about a Magic Treehouse book, that doesn't mean they are able to read for the purposes of a university course of any type.
And no, we do not need to guess at literacy performance of people of the past, we have the curricula, we have the books, we have the essays and exams. There is no "guessing at the performance of imaginary elites from the past", we can use the same texts and the same tests. That's one of the neat things about reading, you don't have to guess what happened in the past, you can just read it.
If you need to apologize for defending Aristotle after more then 2000 years after his death, you must be a part of a political establishment which is very out of tune with the feelings of ordinary people and are a part of a religious community which is unforgiving of any transgressions and can cancel and outcast you for a minor verbal delict.
>>And, are you now signalling yours?
Not very precisely. I think at least 90% of political body would disagree with their position. Also all the biggest world religions believe in repentance.
Claude Design with Fable 5 is an absolute killer app that you’d have to pry from my cold dead hands.
I’m not an Anthropic fanboy - Codex has been my daily coding driver for the last year.
But my point is these companies are building app + model combos that are very sticky. Certainly not as much as an OS, but much more than the “there is no moat” crowd give them credit for.
There are soooo many clones of Adobe Photoshop including clones that run right in your browser, and yet people continue to pay Adobe hundreds of dollars a month because there's like one tiny tweak that Photoshop has, or some small mouse gesture that the clone doesn't completely replicate.
I'm almost on board with what you're saying, except Photoshop had a two-decade advantage both in its technological advancement and in becoming the industry standard, including by infiltrating schools and being pirated by students who would become professionals. American AI companies don't have that gulf.
This is peanuts compared to a major cybersecurity catastrophe that’s surely in the making.
To give credit to the technology and the people using it - and I’m not being facetious - it’s actually incredible that at the current levels of usage the unprecedented catastrophic event has not yet happened.
some things never change. Pre AI I was always shocked that such large and complex systems actually run as well as they do. Especially after getting to see how the sausage is made/works.
It was the mid 2010s when I sensed a lot of SaaS becoming popular. Just host your ticketing systems, your IT management planes, your security management consoles, your SOC, all off-premises.
I wonder if businesses are thinking of ever swinging back to locally hosted, with the increased hostility of the Internet re: AI, vulnerabilities, DoS, and so on.
I'm sure some businesses are considering moving back to on-prem, but for many, I suspect the cost to find onboard, and pay the SMEs to keep those systems running well enough to not fail due to one reason or another isn't as appetizing to them as the ability to offload that work, along with the legal responsibility.
When something goes wrong, pointing the finger at someone else is far easier for most than pointing it at yourself.
One thing that you need to understand is that the usual business manager absolutely hates depending on technical expertise, and that the modern corporate world is fanatically anti-intellectual.
Vendor lock-in? compliance and security risks? stupid systems that cost the company an arm and a leg? nobody fucking cares.
Now, depending on an 130 IQ Engineer that basically holds the whole enterprise on his head? Anathema!!!!!!! Bus Factor!!!!
Also, the 130 IQ Engineer usually can't afford to wine and dine the business manager the way Amazon can. Or dangle the possibility of a cushy yet prestigious sounding job in front of them.
Any individual human (or frog, obviously) is getting out of the pot when it gets uncomfortable.
True stupidity requires a group of humans, all sitting in the pot, telling each other how lucky and special they are to have this wonderful pot, getting paranoid about outsiders who might disrupt their god-given pot-dwelling way of life, and mocking anyone who suggests that the pot might be getting a little too warm.
Clearing LLMs out of our business infrastructure is going to be a massive undertaking. Though I have a tech background, I work in commercial real estate. We are recently seeing new levels of idiocy from the employees, including real estate brokers with zero tech knowledge "coding" solutions to find sites for clients and blindly trusting the output (which I came to find out was complete bullshit), as well as some who have literally stopped communicating with any of their own language - meaning every interaction they have with anyone not in person is made by an LLM. It's a massive threat to our brand and has got to stop. I can't imagine what companies with thousands or tens of thousands of employees who have really been riding the LLM train are going to have to deal with. This thing is more of a virus that exploits human laziness than actual useful tech.
I'd be scared shitless to even try something like this. There is just a pretty website, a video, and a blog post. No info on the founders, I can't find anything on LinkedIn, just a company Vineyard Finance LTD that was incorporated last year.
We're all unhinged about the data we're giving LLMs but here I'd draw the line. I'd rather keep paying the small amount I pay to have my accounts done.
(I'm not factoring the benchmarks into the discussion, because I've never quite cared about them)
reply