Hacker Newsnew | past | comments | ask | show | jobs | submit | hyperpape's commentslogin

You're half right, half not.

This is evidence, but it would be a mistake to think that high profile online contests would have the same dynamics as low profile online polling. 4chan is characterized by an insane level of effort on the specific jokes that they care about. But they don't care about a random poll nearly as much as they care about big stunts like this.


Of course, but a low profile poll also needs less effort to manipulate, so it scales down for both ends.

You’re only counting JEPs, which are only for more involved features. There are lots of changes to the JDK apis that are used by other runtimes. See, for instance: https://javaalmanac.io/jdk/27/apidiff/26/.

Admittedly, the terminology here is almost designed to be maximally confusing, and I’ve never read a good post that laid out how everything relates.


Fair, although even there I don't know if I'd call that "lots" at just 23 added or modified methods that aren't in preview

There are also bug-fixes and performance improvements that are not going to show on the page I linked.

I do think it’s plausible this is a smaller release. Not that this was the real point of the discussion, but I think it’s still just a good idea to have more than one release a year. It keeps things moving smoothly, and lowers the cost of missing a release, which has beneficial effects.


I started programming when Java 6 was relatively new. Back then it was about 3-4 years in between releases. Although I don’t write much Java any more, I’m happy to see changes shipping more frequently now.

Crossing the street and Russian roulette both have non-deterministic risks of injury. And yet I would be bothered to find out that that on my way to work, I was playing Russian roulette by surprise.

Yes and if you decide to play not using a game rulebook but a website that call it "game A" you can't be shocked if "game A" switched from one to the other, even though the doc said opposite yesterday.

That would absolutely be a dick move on the part of the website and would confuse users.

Also, in this case, the game name is not “Game A” but something like “Deep Seek v4 Pro”, which they have previously chosen to use to describe Deep Seek v4 Pro, not Deep Seek v4.1 Flash.


You should look at why you are protecting your earlier comment rather than changing your mind or understanding why so many people disagree with you.

I’ve seen several false (or apparently false) accusations of LLM authorship on HN/Lobsters.

However, we have to distinguish a few hypotheses:

1. No careful readers will notice when a piece is AI written.

2. Careful readers will generally not notice AI writing.

3. Everyone who writes comments on HN will reliably classify writing as AI or not.

Yes, 3 is not true, but Bryan’s point depends on something in the area of 2.

The ability to distinguish AI writing depends on having a good ear. For people who lack it, they either don’t notice and don’t care, or they make paranoid accusations against anything that is remotely non-standard (“you used an em-dash, you must be AI!”).


I'm genuinely curious about the effect, but I simply don't have patience for the AI writing. Can anyone give an actual non-garbage explanation with some respect for the reader?

Slightly less annoying summary from ChatGPT free: https://chatgpt.com/share/6a9ac7a3-15a0-83eb-8c2a-6f72cd9beb....

Caveat emptor: it makes high level sense, but I haven’t thought about it in detail.


I'm genuinely curious: I read through the article and the narrative felt deliberate and didn't trigger my AI-radar. What made you think it was AI written?


It’s rife with “and that’s the important thing. $thing”-type constructions. Human writing uses those much more sparingly and doesn’t separate them into two sentences as often. It has a lot of “subject change: teaser” constructions, too. And overuse of bold and restatement.

Could it be human? Sure. But it doesn’t seem likely to me.


It didn’t read AI generated to me. I also love the irony of then using AI to generate a summary


> This article is the story of chasing that number down to a single machine instruction, and then finding out that the instruction was only half of the answer.

> So the difference has to be in what the JIT generated, and the profiler gives us exactly that.

> That is the whole vocabulary. Let’s read some code.

> Decoding the x86 version instruction by instruction is out of scope here.

> What is not architecture specific is the logic.

Here’s a segment flagged by Pangram: https://www.pangram.com/history/87e25169-30a4-4030-a65a-dba8...


I don't know if it's AI or not, but if it is, it isn't egregious, and certainly not "garbage" or disrespectful to the reader.


I will admit I shouldn’t have said “garbage”, because it is too emotional.

But I absolutely think it is bad for the reader, and deserves to be called out. Authors should know that it’s not good enough.


Even if that were AI (and I think there are better examples in the article to indicate that it is), why do you care if you just used AI anyway to summarize it anyway?


I don’t really care that it’s AI, the problem is that it’s bad AI writing.

The summary I shared is much more straightforward. It’s mediocre, but just barely good enough to extract the message without making me super annoyed.


I don't think the Pangram result is shareable to others.

"If you are sure the link is correct, please make sure you are logged in with the account associated with this result."


If you don't like AI writing, why are you on hacker news? Most of the articles posted are written by AI.


Some of us have been here longer than AI. "To be intellectually stimulated" would be the why, which there used to be more of prior to the frontpage becoming riddled with AI slop. "HN is for conversation between humans" as the guidelines say, and though the guidelines imply that's towards comments, I'd prefer it to be towards submissions too, for the same reasons it was added for comments.

I'm already at the point of "the reward/effort of HN is getting pretty low", but the question is, where to leave for?


Don't think everything is just "who can produce the biggest/smallest number": https://mathstodon.xyz/@tao/117208619314517025.


You're right that Shin Jinseo is a generational talent, and more dominant than anyone since Lee Changho (peaked in the 90s and was strong into the early-mid 2000s).

However, you can't compare goratings over time, the top ranks are not nearly stable enough. https://www.goratings.org/en/history/ (I think it's believable Shin Jinseo is better than Lee Changho, but not that there has been steady progress since the days of Lee Changho, so that there are now 20 players stronger than him).


You can't compare Elo ratings over long stretches of time, period.

Ratings drift over time, based on the total population of people competing. I think the most accurate way to view Elo ratings is as a measure of skill vs. the average rated player.

If you want to compare Magnus Carlsen's peak rating of 2882 in the year 2014 to Garry Kasparov's peak rating of 2851 in the year 1999, you have to know how strong the total pool of players (including all the amateurs who compete at lower levels) was in the year 2014 vs. 1999.

The only way to actually anchor the Elo system over long periods of time would be to have rated humans occasionally play against a set of unchanging computer players, which could then serve as static rating calibrators. You could use those games to then calibrate Elo ratings from different time periods to a common scale, by asserting that the computer ratings don't change.


Even that would be slightly malinformative because of opening theory. If you warped a very strong player from the past to the present, he'd do very poorly at first simply because of advances in opening theory. But give him a bit to catchup and he'd likely have his rating zoom on up. So modern players would do better against the static computer because of the same advantage, but that doesn't mean they're necessarily stronger in the sense that we hope to measure. The question people always want to know are things like how would a Morphy, Capablanca, or Alekhine do in modern times with access to modern theory and the like - not how well would they do against Carlsen if they went in with nothing but the knowledge of their era.


In Go, we had an extensive documentary on this, based on the strange case of a man possessed by the spirit of an ancient master Go player. I'm not going to spoil the results here; it's an enjoyable series and includes many small introductory tutorials to the game.

https://en.wikipedia.org/wiki/Hikaru_no_Go


> You can't compare Elo ratings over long stretches of time, period.

I agree that you cannot do it reliably, in principle.

However, Bobby Fischer peaked at 2785, and Gary Kasparov peaked at 2851. These are not far from what informed observers suggest--maybe 50 or 100 points off. They are well into super-grandmaster territory. Kasparov would be 1st today, Fischer would be 4th.

But on goratings.org, the top player of the 90s would be roughly 30th today.

My point is that the go ratings are much more unstable than the chess ratings. With chess ratings, you'll be wrong in the details. With go ratings, you'll be catastrophically wrong.


I agree that the drift in chess Elo ratings is slow enough that we can say Bobby Fischer at his peak would still be a strong grandmaster today (if he were given some time to study developments in opening theory). But I don't think we can say whether he would land at 2700 or 2800 in today's Elo scale.


This is within the realm of reasonable opinion, but I'd tentatively say it's an uncommon one--I thought it's generally agreed that the very best players of that era didn't lag contemporary players in terms of skill, only theory. I may be wrong, though.


Chess has developed significantly over the last 50 years. The best players now can be reasonably expected to be much more skilled than the best players of 50 years ago, even if you discount opening theory knowledge.


> If you want to compare Magnus Carlsen's peak rating of 2882 in the year 2014 to Garry Kasparov's peak rating of 2851 in the year 1999, you have to know how strong the total pool of players (including all the amateurs who compete at lower levels) was in the year 2014 vs. 1999.

I imagine that FIDE has all of this data somewhere, right? They don't publish more in-depth distributional analyses of players somewhere?


The distribution of Elo ratings will not tell you how strong the average player is in an absolute sense.

Elo ratings measure differences in skill between different players. A player rated 400 points above another player will win 90% of the time. Only rating differences are meaningful. Absolute ratings aren't.


Yes, I can understand why absolute ratings are impossible to extract.

However, since the population maintains some continuity over time (players gradually enter and then leave over time), would it not be possible to reconstruct relative ratings between players that didn't play during the same era?


Only if you were to assume that a player's prowess remains constant throughout their career, which we generally know to be false. (I'm completely inventing dates here) If Fisher played Kasparov in 1990 and Kasparov played Carlsen in 2020, you can only compare Carlsen to Fisher if you assume Kasparov's skill was about the same for this entire duration, which no one believes to be the case.


A lot of people don't realize that chess players peak around the same age as athletes do. Carlsen was at peak dominance in his mid-20s. Kasparov peaked later, at 36 years old, but still, it wasn't at age 50 or 60.

Chess requires you to be really sharp. Experience increases with age, but there's some cross-over point at which the decline in calculating speed is more important than increasing experience. Heck, I'm measurably better at chess tactics in the morning after a good night's sleep than in the evening after work.


I'm sure it's also a matter of how much time you actually dedicate daily to this - and it's quite hard to maintain passion and focus for 8+h /day on the same game for 30-40 years.


The problem is ambient go knowledge. A top 100 player would easily beat time traveling Lee Changho in his first few matchups. Of course give peak Lee Changho a fortnight to prep with Katago and … well that would be something!


It's not obvious to me, but I lean towards saying this is false. The 2026 player would play a better AI inspired opening, but I'm not sure that would be enough to overcome the skill difference.

(This is especially true if they don't get briefed "this is time-traveling Lee Changho, he doesn't know contemporary joseki, play a trap variation").


Your point is pretty fair.

The reason I think the top 100 pro would win is that when I watch pros review they are very quick to catch when someone is playing a non-current or up to date style. I’ve also seen a few middle game situations where the classic famous go move is now “obviously” not good after we’ve seen how AI handles it.

This all adds up to many ways good sound moves to Lee Changho have now obvious responses. Do I believe he would adapt inhumanly fast? Absolutely! Legendary fighting spirit


Could you just have superhuman Go AI just how good humans are somewhat more objectively?

Not without flaws of course, but probably interesting


It significantly predates Claude, and has been one of, if not the best engine in the world for many years.


If you scaled Shin Jinseo to 2800, you would have players with extremely negative ratings. This page shows ratings of European players on a roughly aligned scale: https://europeangodatabase.eu/EGD/createalleuro3.php?country.... It still has negative numbers on it, and this only contains players who have attended a tournament (though it's more common for beginners to play tournaments in the west, since it's hard to find times to play).

It's not a comparison of the worth of the games (I play both, though I'm better at Go, and prefer it), but the dynamic range of Go is larger.

That said, any cross-game/sport comparisons of this kind are pretty tough to do properly.


This is a bit of a digression, but: It's an interesting question what (if anything) that larger dynamic range means.

One thing you'll hear people say sometimes -- I've said it myself -- is that this shows that in some sense go is a "deeper" game than chess; there's more to know and understand, more variety of possible human skill.

That might well be true. It certainly feels a more elegant game, and involves longer tactical sequences, and so forth. But this may be misleading.

Consider the game of treblechess. To play a game of treblechess, you play three games of ordinary chess and look at the overall result.

Suppose that when we play ordinary chess, I win with probability W, lose with probability L, and draw with probability D = 1-(W+L). And suppose separate games are independent of one another (which might not be true in reality, but never mind). What happens when we play treblechess?

I win 3/0 with probability W^3. I win 2.5/0.5 with probability 3W^2D, because that happens when I draw any one of our three games and win both of the other two. I win 2/1 with probability 3(W^2L+WD^2), because that happens when I win two and lose one or win one and draw two, and for each of those there are three choices for which game is which. So I win at treblechess with probability W^3 + 3(W^2(1-W)+WD^2).

I can draw by getting one each of WDL (probability 6WDL) or by drawing all three (probability D^3).

Suppose that when we play chess I win 40% of the time, draw 50% of the time, and lose 10% of the time. Then our Elo difference is about 107 points. In triplechess, I will win 65.2% of the time, draw 24.5% of the time, and lose 10.3% of the time. Our Elo difference is about 214 points.

If in ordinary chess I win 65% of the time, draw 25% of the time and lose 10% of the time -- about the same odds as for treblechess in the last example -- then our chess Elo difference is about 215 points. At treblechess I will win 84% of the time, draw 11% of the time, and lose 5% of the time, and our Elo difference will be about 375 points.

If in ordinary chess I win 15%, draw 75%, lose 10%, then our Elo difference is about 17 points; in treblechess I will win 31%, draw 49%, lose 20% and our Elo difference will be about 41 points.

Treblechess Elo differences are on the order of double ordinary chess Elo differences! Clearly treblechess is a game with twice the depth of ordinary chess!

But it isn't. It's just longer and gives more opportunities for the better player to come out ahead overall.

Go is also a longer game than chess, though of course not in the same way as treblechess is. Perhaps the larger Elo range of go is more because of that than it is because of actual deeper strategy and tactics?


As an amateur (but competent) player of both, when I play go it almost feels like I'm playing multiple smaller games at the same time on the same board. Even if I might be struggling in one area, I can make it up in another. So it's very hard for a lesser opponent to beat me, even if they do gain an advantage in one area. Likewise in reverse when I'm playing a superior opponent. I might get in a nice kill, but next thing I know we're playing a different battle on the other side of the board and I get crushed.

So I think there's truth to what you're saying.


Maybe the depth of a game is related to whether it will scale. Go is played on different board sizes and still works.

If you make a Backgammon board bigger it would just be a slog and no real increase in tactical challenge.

Chess rules dictate a set size of board, which I guess has been refined over time.

If a Go board is made bigger or smaller it just adjusts the problem space, the rules and core of the game remain the same.


Perhaps interesting: Japanese chess, Shogi, is usually played on a 9x9 board, but there are many variants, which are played on bigger boards. Though games take longer and those variants are played rarely. It scales, but I think it's fair to state, that Go scales much better, due to its simplicity (simple != easy).


I’ve thought about this before and come to very similar conclusions. The ELO range between the best and worst players is meaningful, but you can’t just read off the depth of the game from it in the senses we most care about.


I'm unfamiliar with Go ratings, but chess ratings are based on the Elo system which is a simple mathematical prediction system. Borrowing some figures from Wiki [1] we get:

  1.00 +800
  0.99 +677
  0.9 +366
  0.8 +240
  0.7 +149
  0.6 +72
  0.5 0
  0.4 −72
  0.3 −149
  0.2 −240
  0.1 −366
  0.01 −677
  0.00 −800
The second column is your rating minus your opponent's, and the left is your predicted result. So if you are rating 1849 and your opponent is rated 1700 then you'd be expected to score about 70%. To have a 1% expected score against Magnus, you'd need a rating of about 2150.

[1] - https://en.wikipedia.org/wiki/Elo_rating_system


I was under the impression that Go also uses Elo, then I did a bit of cursory research and discovers that it varies.

Two major federations are American Go Association (AGA) and European Go Federation (EGF). EGF uses an Elo-inspired update rule since 2021. AGA uses a quite-different Bayesian system without pairwise update; they provide a paper and a C++ reference impl.

Asian countries don't bother with such numeric ratings. Instead, rankings are titles which are won through tournament promotion structures (sounds similar to Sumo to me).

Interesting, because I always thought that it was more "apples to apples", and that the higher upper limits of Go rankings was somehow indicative of the higher "dynamic range" of the game compared to chess. For example, if Elo were applied to basketball, what would the Elo of Lebron James be compared to a playground hooper (leaving aside that 1-on-1 isn't the best part of Lebron's game)... would it be higher or lower than Magnus Carlsen in chess? I don't have an intuition.


While it's not official, GoRatings (https://www.goratings.org/en/) ranks all players based on a variation of the Elo system (note that it is international, even if in practice you'd have to scroll quite far to find a few lonely Western persons).


AGA and EGF are very minor federations in the Go world, all the professionals are in Asia where the game is far bigger.


It works the same in go. Just, as the parent has said, the dynamic range of go is higher.


Maybe, but that's sort of begging the question that those open weight models aren't significantly trained using "distillation"[0]

[0] not technically distillation. https://thomasdullien.github.io/posts/2026-06-15-rl-economic...


distillation is a minor piece of training data, you have to have a good foundation for it to be helpful, and even if you have good traces, you need a good RL reward scheme at the point it is used (very challenging)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: