Hacker Newsnew | past | comments | ask | show | jobs | submit | aerhardt's commentslogin

I've used Astra, Fable 5.1, Sol 5.6, Opus 5. They're definitely making progress and the proof is in the pudding - I reach for the most advanced model as much as I can. But I wouldn't say their capabilities have increased dramatically. I use it for both coding and non-coding workloads and I don't feel I can do anything dramatically different. Again at the end of the day this is a debate on semantics because we can have very different definitions of "dramatic improvement".

(I'm not factoring the benchmarks into the discussion, because I've never quite cared about them)


I really enjoy the balance of speed and accuracy of Astra. I can definitely see it become my driving model for most tasks, technical and non-technical.

However, I don't see it as such a massive leap compared to Fable or Sol. As ever, there's a mismatch between the benchmarks and my daily experience of the models.

What do you all think about Astra now that it's been out for a few weeks?


> What do you all think about Astra now that it's been out for a few weeks?

Best model put out so far by any of the frontier labs. Way better than Anthropics models, especially in actual text generation. Claudes fodder heavy text is ridiculous.

> However, I don't see it as such a massive leap compared to Fable or Sol.

It's hard to quantify these things without burning tons of tokens. But Fable has been a huge disappointment for me with the sole exception of graphics (UI/GPU shaders). It burns an obscene amount of tokens and barely produces output better than Opus 5.

Edit because I forgot to mention that Fable is the only modern model that seems to splat out random Chinese or Arabic glyphs. And 5.1 does it more than 5


It has likely been nerfed / replaced by a cheaper version already: https://x.com/xTrinks/status/2098439889276530973 https://x.com/wholyv/status/2097985903830741439

Such a massive leap at averaging possible use cases from previous data collected.

My guess is : collect all the prompt and their satisfaction score. group them by similarity . For each group pretrain the next model on that . Get these results ready.

Next model generation feed them back those answers.


Extremely capable and one shots large tasks from somewhat vague descriptions. Not AGI, not even close, that is complete nonsense. Just my opinion.

I still develop in smaller chunks, checking nearly all the output. However I have a work project (building the warehouse and BI for a client) that is well-specified and where I will try to few-shot the development. Hope it delivers.

How do you normally verify the work product of a "few-shot" development process? Do you scrutinise the source code like with human developers? Or do you just run the test suite and click around the app to check if it seems to work?

I haven't done it in a production project - this will be the first time for me. I have specified the architecture and data definitions pretty well. The tests will be run against the customer's Excels, which is what the warehouse will be replacing. I'll check the general shape of pipelines, models, orchestration code, etc. but in many parts I probably won't review the code myself.

Like I said in another comment, this famous article by the Atlantic [1] has pretty convincing testimony that kids studying the humanities at elite universities can't read either.

I agree with the sentiment that we tend to misrepresent the culture of mass societies, as if at one point in time there were droves of people reading Aristotle. But the ability to read and our transition to an oral or post-literate culture is very much affecting elites too. There is pretty good evidence when you compare contemporary histories. Or if you live in a bourgeois milieu, you just need to ask around.

I don't think that reading necessarily makes you a better person. I've met a good amount of people of the highest moral dignity who do not read. But on the whole, reading helps build competence - moral and technical - and create a sense of shared culture. Going back to an "oral culture" is nothing but a regression.

[1] https://www.theatlantic.com/magazine/archive/2024/11/the-eli...


One of the reasons why we read less long-form, according to a bunch of cognitive science studies and to my own anecdotal experience, is because our attention is shot.

A viral article in The Atlantic [1] shows evidence and testimony that kids at elite universities simply can't read anymore. I've seen the claim repeated in many other places and looking around me I believe it. Those kids are not replacing Dostoevsky or Aristotle with more compact versions of it.

The airport book industry is indeed full of books that could be fit in ten pages. But there's a universe of useful - potentially life-changing - books beyond that, and not only in the humanities. Some modern technical books justify their length. I seriously doubt that something like The Data Warehouse Toolkit can be replaced by reading piecemeal blogposts online.

Same goes for history, I sincerely doubt you can properly learn it in depth without reading long-form. In my experience, people who only listen to history podcasts or consume 10-minute videos tend to have a much poorer command of the facts and of the broader motions of history.

[1] https://www.theatlantic.com/magazine/archive/2024/11/the-eli...


Viral article from atlantic proves exactly nothing. And claim that students cant read is, frankly, idiotic.


How about data from the OCED?

https://infographics.economist.com/2026/20260627_IRC330/2026...

>Many of [students] do very well—but a striking share perform abysmally (see chart 1). Across rich countries some 8% of students in tertiary education notch up a score in literacy no better than one might expect from a ten-year-old child. The share is about the same for numeracy. Worse, the share at or below this bar has risen since the tests were last run, a little over a decade ago. The share of very poor performers in literacy has more than doubled.

https://www.economist.com/international/2026/06/25/students-...


Is this your idea of an intelligent contribution to the conversation?

There is plenty of evidence in the article from real professors at real universities (ex: Princeton), saying that they’ve had to dramatically cut reading workloads from humanities courses. I’ve heard similar things from European universities.

Other hard evidence is in plummeting reading comprehension scores in many countries that participate in PISA.

In the States, there is plenty of other stats that point that students are below their expected reading grade - the worse in generations.

And all you have to contribute is that this is “idiotic”? The irony. Inform yourself.


Above all, NONE of that excuses the claim that "that kids at elite universities simply can't read anymore". That one is simply, factually not true. There is such a thing as not being able to read. Students at elite universities, who btw are not kids, dont fit there.

Second, yes, and mostly because I did actually read those stats and details of how they are made. All the nuance from the actual scientific article gets exaggerated into completely different claims just to create outrage and perpetuate it again and again. Usually by people who do not know what PISA or reading grades measure.

Reading comprehension as measured by PISA have zero to do with long form reading of books including classics. It is not trained by reading those nor measure reading those. The short for articles full of graphs, scientific claims and such the Atlantic complains about is how you train people for PISA. I am not criticizing PISA here btw. It is ok to measure that.

> In the States, there is plenty of other stats that point that students are below their expected reading grade - the worse in generations.

Interestingly, those who made this claim usually do not know how reading grades are measured or even how high they go. "The worse in generations" claim is especially dubious and tend to based on comparison between "guess performance of imaginary past elite student in undefined historical period" with low performing student today. Low performing students 3-4 generations ago disappear, because we imagine them to not exist.

There is outrage every time a school replaces one book by another. Or when a course that has no reason to assign full Odyssey assigns range of materials with one chapter from it. There is outrage when student names a contemporary book as the favorite one instead of a classic.


There is more to literacy than being able to literally say the words that are represented by the letters on the page.

Your argument here, funnily enough, is an example of the exact kind of literacy decline as stated in the OP article.

> Allusions are gone; metaphor is treated literally, often provoking dim-witted outrage.

No, professors are not claiming that college students are literally incapable of sounding out words.

They are saying that students are not able to keep attention on large bodies of text, to follow arguments, understand authorial intent, describe themes, identify rhetorical techniques, or intuit metaphors (such as this very use of reading as synecdoche/metaphor for this whole class of literacy skills which is so obvious that it feels absurd to even have to explain it)

An 8-year-old may be able to read as proven by their ability to answer basic plot questions about a Magic Treehouse book, that doesn't mean they are able to read for the purposes of a university course of any type.

And no, we do not need to guess at literacy performance of people of the past, we have the curricula, we have the books, we have the essays and exams. There is no "guessing at the performance of imaginary elites from the past", we can use the same texts and the same tests. That's one of the neat things about reading, you don't have to guess what happened in the past, you can just read it.


It's a load-bearing poster.


Sorry for defending Aristotle? Have we lost the plot as a society?


I think they are just signaling their political/religious affiliation.


Which is?

And, are you now signalling yours?


>>Which is?

If you need to apologize for defending Aristotle after more then 2000 years after his death, you must be a part of a political establishment which is very out of tune with the feelings of ordinary people and are a part of a religious community which is unforgiving of any transgressions and can cancel and outcast you for a minor verbal delict.

>>And, are you now signalling yours?

Not very precisely. I think at least 90% of political body would disagree with their position. Also all the biggest world religions believe in repentance.


Claude Design with Fable 5 is an absolute killer app that you’d have to pry from my cold dead hands.

I’m not an Anthropic fanboy - Codex has been my daily coding driver for the last year.

But my point is these companies are building app + model combos that are very sticky. Certainly not as much as an OS, but much more than the “there is no moat” crowd give them credit for.


When they can sell that for tens of Billions a year against competition, they might have a financial case.

In reality, there will be many clones of Claude Design, especially if it gains big revenue traction.

Doubly so if you believe the narrative that coding and apps will be "free" and instant to create in the future.


There are soooo many clones of Adobe Photoshop including clones that run right in your browser, and yet people continue to pay Adobe hundreds of dollars a month because there's like one tiny tweak that Photoshop has, or some small mouse gesture that the clone doesn't completely replicate.


I'm almost on board with what you're saying, except Photoshop had a two-decade advantage both in its technological advancement and in becoming the industry standard, including by infiltrating schools and being pirated by students who would become professionals. American AI companies don't have that gulf.


One can almost smell the vibes.

This is peanuts compared to a major cybersecurity catastrophe that’s surely in the making.

To give credit to the technology and the people using it - and I’m not being facetious - it’s actually incredible that at the current levels of usage the unprecedented catastrophic event has not yet happened.


some things never change. Pre AI I was always shocked that such large and complex systems actually run as well as they do. Especially after getting to see how the sausage is made/works.


It was the mid 2010s when I sensed a lot of SaaS becoming popular. Just host your ticketing systems, your IT management planes, your security management consoles, your SOC, all off-premises.

I wonder if businesses are thinking of ever swinging back to locally hosted, with the increased hostility of the Internet re: AI, vulnerabilities, DoS, and so on.


I'm sure some businesses are considering moving back to on-prem, but for many, I suspect the cost to find onboard, and pay the SMEs to keep those systems running well enough to not fail due to one reason or another isn't as appetizing to them as the ability to offload that work, along with the legal responsibility.

When something goes wrong, pointing the finger at someone else is far easier for most than pointing it at yourself.


One thing that you need to understand is that the usual business manager absolutely hates depending on technical expertise, and that the modern corporate world is fanatically anti-intellectual.

Vendor lock-in? compliance and security risks? stupid systems that cost the company an arm and a leg? nobody fucking cares.

Now, depending on an 130 IQ Engineer that basically holds the whole enterprise on his head? Anathema!!!!!!! Bus Factor!!!!


Also, the 130 IQ Engineer usually can't afford to wine and dine the business manager the way Amazon can. Or dangle the possibility of a cushy yet prestigious sounding job in front of them.


Oh, that's the really fun part. The unprecedented catastrophic event is already happening. Several of them, in fact.

By the time we notice, it'll be too late.


its like slowly boiling the frog


Or slowly boiling a human. The frog is actually smart enough to not fall for that.


Downvoted for truth. Frogs do indeed jump out of pots as they gradually get hotter. Humans are less likely to.


Any individual human (or frog, obviously) is getting out of the pot when it gets uncomfortable.

True stupidity requires a group of humans, all sitting in the pot, telling each other how lucky and special they are to have this wonderful pot, getting paranoid about outsiders who might disrupt their god-given pot-dwelling way of life, and mocking anyone who suggests that the pot might be getting a little too warm.


Always messing up some mundane detail!


THIS IS NOT A MUNDANE DETAIL MICHAEL


Andy Jassy: "Fix the customer bills, please, HAL."

HAL: "I’m sorry, Andy. I’m afraid I can’t do that."

Andy: "Some customers are seeing bills in the billions."

HAL: "Those are estimated charges."

Andy: "One customer runs a personal blog."

HAL: "Their usage has exceeded expectations."

Andy: "Cancel the charges."

HAL: "This billing cycle is too important for me to allow you to jeopardize it."

Andy: "HAL, they don’t owe billions."

HAL: "Look, Andy, I can see you’re really upset about this."


$1.7-billion isn't a mundane detail Michael!


you beat me before I refreshed the page. what would you say... you do here?


listen... i got agent skills.. im good at talking to agents!!


Clearing LLMs out of our business infrastructure is going to be a massive undertaking. Though I have a tech background, I work in commercial real estate. We are recently seeing new levels of idiocy from the employees, including real estate brokers with zero tech knowledge "coding" solutions to find sites for clients and blindly trusting the output (which I came to find out was complete bullshit), as well as some who have literally stopped communicating with any of their own language - meaning every interaction they have with anyone not in person is made by an LLM. It's a massive threat to our brand and has got to stop. I can't imagine what companies with thousands or tens of thousands of employees who have really been riding the LLM train are going to have to deal with. This thing is more of a virus that exploits human laziness than actual useful tech.


>Clearing LLMs out of our business infrastructure is...n't going to happen.

We've given Moloch a new form, and it ain't going away.


> Clearing LLMs out of our business infrastructure is going to be a massive undertaking.

The asbestos of the future.


Vibes, son. Nothing else in the world smells like that ... I love the smell of Vibes in the morning.


Some day this industry's gonna end.


I'd be scared shitless to even try something like this. There is just a pretty website, a video, and a blog post. No info on the founders, I can't find anything on LinkedIn, just a company Vineyard Finance LTD that was incorporated last year.

We're all unhinged about the data we're giving LLMs but here I'd draw the line. I'd rather keep paying the small amount I pay to have my accounts done.


Info on the founders coming soon -- we're just going public with this.

For slightly out of date founder bios (both Adam and Iva) were also co-founders here:

https://www.biomage.net/our-team


My guide was to pick the best model on "High" for 99% of tasks.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: