Hacker Newsnew | past | comments | ask | show | jobs | submit | jesse_dot_id's commentslogin

The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.

AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.


I doubt that would change the perception. Every model release is followed by accusations of nerfing.

There are several projects that repeat benchmarks on published models. None has ever found significant fluctations

Here's one example https://marginlab.ai/trackers/claude-code/

Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.

This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.


This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.

If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.


> This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.

Click the part at the end that says "View historical performance". They wait to collect more data about a new model before adding it to the overall charts.

The overall solution rate continues to climb when new models are considered.

> If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.

The y-axis is amplified to make differences look larger than they are.

Hover over the dots to see the confidence interval. A 1-2% change means nothing.


The page is currently tracking Opus 5, which was released July 24 and has not seen any updates since.

30-day average 83% [71-91], last result is 79% [66-88], and it dipped to 75% [61-85] a week ago.


The page says

> We always use the latest available Claude Code release and the SOTA model

They've stepped up the benchmark for each new model. Presumably they're gathering Fable data now.

They also link to Anthropic's public blog post about some degradations, their cause, and how they fixed them. The time period sounds like the "recent weeks" you experienced: https://www.anthropic.com/engineering/a-postmortem-of-three-...


Fable 5.1 was released 20 days ago. That's a lot of time to gather data. Since they're still tracking Opus 5 anyway, there is no obvious reason the perf delta section would be disabled. Are you involved with the project?

The Anthropic post points to the latest fix on Sept 12, and the issues mentioned also only affected Sonnet 4 and Haiku. Opus was misbehaving just last week. You are choosing to not see the evidence of degration, 85% -> 75% is a generational dip in intelligence.


Not involved, no. Just guessing. It doesn't look frequently updated.

A lot of these daily-benchmarking sites popped up earlier this year. Most of them have faded away after they all failed to produce the smoking gun that everyone expected. This site survived because it kind of caught a dip one day, maybe.

The results are really rather flat and daily benchmark runs are expensive, so most of these projects give up after a while.


It would change my perception but only if there were a competent and stringent administration in place. I didn't used to have to wonder if the ground beef I was buying was actually 1lb because there were inspections and repercussions, but stuff is kind of chronically underweight these days.

A properly run OWM enables you to stop wondering if you're being ripped off and that's what AI needs because I think it's incredibly easy to just assume we're being ripped off because these companies are all built on a foundation of wonton theft. (Not that I really care about that — I think all information should be free, but still.)


If you go to https://marginlab.ai/trackers/claude-code-historical-perform... there is a very clear downwards trend in the two weeks before Opus 4.7 release. Then a sudden and dramatic drop seven days before Opus 4.8. And now we seem to have entered another decline in the last ten days, beyond the usual noise of Opus 5 scores

Anthropic terms of service:

> 12. General terms

> Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services.

> Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them.

You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.


You are not buying something and expecting it to be what's on the tin? Aka what the benchmarks show?

My comment is what is on the tin. As a consumer, yeah, I find this annoying. But do I want to bring the full force of government regulation on it? That's quite a strong reaction

Well yes, it's no different from breaching an SLA agreement.

THIS EXACTLY.

The only regulation that we need right now is the model that's on tap


we could even just repurpose the same office, "weights and measures" is oddly relevant

Petition to rename them to the Office of Weights and Biases, haha.

I thought METR was supposed to fill this role?

... But what exactly is the "weight" metric you have in mind?


Yeah, not touching xAI for several glaring reasons. I share your confusion.

Good for you. Some people paint on canvases, too.

Let's have the nationalization argument with literally any other US administration in place.

yeah me too on that. I can literally hear the bailouts getting stacked right now too haha

It will be interesting to see if this changes because presumably AI is using very predictable historical models, but it seems like the climate is shifting into something unseen that we won't have models for?

Yes, I heard from one first-rate forecaster that he thinks AI forecasters are especially weak in predicting big disruptive changes to the world.

Hard to study this, obviously!


Are you referring specifically to climate as in weather? The article is about forecasting a range of future events, not specifically weather.

Climate as a pattern of weather over a long period of time. If the climate is increasingly unpredictable, I would think that it wouldn't really effect our ability to make short-term predictions, like a few days out.

But our ability to forecast weather on a longer timeline, like for industrial forecasting, is calibrated on historical weather patterns. But with weather being more erratic and unusual, I don't understand how AI will be forecasting with the models they have now.


I think you're misunderstanding my point. The headline's usage of the term "forecasting" is not referring to weather forecasting, or climate forecasting. It's referring to forecasting a wide range of possible future events.

For example, predicting which party will win an election, or if there will be a major cyber event in the next year, or the price of Gold in 6 months. Presumably a few of the questions could be related to climate as you're thinking of.


Of course model predictions will be acted upon, which will invalidate the predictions.

At least 150k on my relatively small FastAPI project, but hit my session limit. Continuing in a few hours.

Oof. YAGNI. 150k tokens is where you start hitting the "dumb zone" (model attention issues and inconsistent adherence to instructions).

Hard agree. The entire tech media is acting insanely gullible in this regard. It's insane.

Always have been. Since at least 10-15y the tech media is pretty much a PR department for the tech industry

Nobody at OpenAI is closely monitoring token usage or egress, eh? Alright. They should potentially fire a whole team of engineers if that's the case. I monitor egress from VMs that don't have shit on them.

No, they're not wrong. There are just a very large number of people who are a special kind of stupid.

Not you though. You're the only smart one, huh?


There are so many other existential risks to humanity. Throw it on the pile. At least this one has a chance to be really cool.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: