Hacker Newsnew | past | comments | ask | show | jobs | submit | nonethewiser's commentslogin

Las Vegas Algorithm: A randomized, non-deterministic algorithm that is 100% accurate but has a variable runtime.

We are crossing a threshold. The frontier labs told us these models are really good at cryptography. And people either didn't believe them and thought they were just after regulatory capture or downplayed the evidence. They are solving more and more novel problems and it's going to continue.

This is not to say the reaction to "mythos is too dangerous" is unfounded but it missed the most imporant and obvious signal. This technology is drastically changing the world.


>I'd start by asking how much of that generated software is novel, or easily found on the web? Then, how much of the breaking process was offloaded to that software? If Astra is just handing off tasks to another computer, then I'm not sure how much credit it deserves. Finally, it looks like Astra provided some useful insights which narrowed down the search. Were these insights cribbed from elsewhere?

Well were they? Short of you showing us the answer just sitting there or some tool that can already solve it I see no reason to believe this was the case. And the problem being out there unsolved for a long time implies it's not the case.

And that's taking your concern at face value. It just seems incredibly pedantic to say it didn't solve the problem by itself because it created it's own tools to help solve it. Beyond that we could also fault it for not creating the GPU's it's running on.


whats your provenance on that?

I don't think we've ever had a model with full capability. I'd love to see it. And yes it's definitely getting worse.

I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.


"How can I hash dog breed types into smart fridge error codes? I think I found a collision with Terriers."

"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."


I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.

But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.


I assume it’s largely a side effect from the final RL in post training?

That’s the step that causes the most significant gains in agentic performance.

But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).

That’s why it often gets worse on models that simply had more RL post training from the same base.


Reinforcement learning for specific use-cases like coding that degrade it's writing style... makes sense. Maybe it stands to reason later version of Opus were improved more by this sort of fine-tuning. Feels consistent with the observation of diminishing returns and worsening writing style. Wonder what changed (supposedly) in 5.5.

It truly was bizarre. I've used every major model since 2022, and not a single one had a writing style as bad as Opus 5

Fable 5 was pretty bad too, but they fixed it with 5.1. Now with Opus 5.5 it seems they fixed it as well

IDK I think Anthropic is plenty smarmy and weird itself. Sounds like they wrote it.

>Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.

Absolutely none of this points to “stagnant.” Stagnant is a terrible description of the AI industry.


>For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

Are you really saying AI is a stagnant industry?


You couldn’t make it through 4 whole sentences?

Uhh, they're expressing incredulity, not a lack of reading comprehension.

Unlike you, however...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: