Hacker Newsnew | past | comments | ask | show | jobs | submit | ngomez's commentslogin

Well, Buckmaster says both his and Alpöge's use of Codex was non-institutional, and OpenAI claims the right to train their models on inputs and outputs of non-enterprise users in their service policies [0]. So I'm not sure they were even promised that.

[0] https://openai.com/policies/how-your-data-is-used-to-improve...


It isn't relevant whether they were promised that. Indeed I think the assumption must be that they were not promised that, since otherwise the author asking if they were would not make much sense.

If OpenAI did use the conversations from Buckmaster and Alpoge, then not disclosing it, explicitly, is plagiarism. If they planned to use that plagiarism to pressure the authors to publish, that is even more unethical. What the terms of use say does not make it any more or less ethical.


If you use their consumer subs, you get subsidized tokens in exchange for them having full access to your data. Those are the T&Cs. Have something secretive? Get a commercial sub with zdr.

(And if memory serves, there is also the opt out from training on consumer subscriptions). Its not plagiarism if you make your data available for the purpose of training their LLMs. It is you giving away your IP for some tokens.


That's absolutely right. Why the downvotes? If OpenAI are using but not acknowledging the work of others that's plagiarism. If they don't know for sure, but aren't performing due dilligence to make sure they aren't, that's also plagiarism.

It's not as simple. All our chats are being used by both labs for their future product (unless signed by ZDR). Where should the acknowledgement begin? Who should be acknowledged? The whole world? All the 2B users of AI?

If I know person A is working on problem B.

I am free to work on problem B too. Why should person A be limited to working on it.


Are you free to intercept person A's emails / hack their computer to find their notes on how they're approaching problem B?

Finding Codex session data in the training set that you tie back to these two researchers is like an hour-long task.

You're also free to plagiarise anyone you want. There are no laws against it on most jurisdictions.

Also brain raping* is not illegal in most jurisdictions.

But they're both deeply disturbing.

_________

* https://youtu.be/JlwwVuSUUfc?si=uWl4-LCHAeI7qtb3


I can imagine excuses for unknowing plagiarism in this case. What is described in the article seems much more serious: a research program that was only initiated following reports of the author's similar program. In this case no excuses of "I didn't know" can apply, it is not like this revealed some obscure work from the 1980s nobody could reasonably have foreseen. And as far as I can tell this program was only really initiated to apply pressure to the researchers, without their knowledge/consent. It looks very weird.

> Why the downvotes?

I think there was only ever one. Not sure why.


Your comment was greyed out when I saw it earlier, maybe you missed some downvotes?

About the plagiarism issue, I model it as OpenAI being an advisor and their AI a PhD student. If the advisor puts their name on a paper behind that of their PhD and it turns out the PhD copied the text of the paper from somewhere else the advisor is also responsible of plagiarism, not just the student. The least the advisor can do is withdraw their authorship from the paper.

But, yeah, point well made: it could be much worse than that. Like an advisor instructing a student to copy someone else's paper.


I think "greyed out" just means "0 points or less", so if you get 1 downvote without any upvotes it'll be greyed out. For instance your initial reply to me is now greyed out, and I have since observed a few upvotes and downvotes on my original comment (the downvotes apparently from people who aren't willing/able to justify why).

Personally I don't like thinking of LLMs like a PhD student, because most PhD students remember where they learned things from, while LLMs essentially cannot. I think of it a bit more like someone using a search tool carelessly. Although in this case it is apparently more like deliberate misuse than carelessness.


OpenAI is willing to credit so plagiarism is not the right framing here.

It's not clear to me that they would have credited the authors if they had not got in touch with OpenAI first.

Only for ChatGPT, if the user hasn’t opted out. Would mathematicians be using ChatGPT for this kind of work? Genuinely asking, I know nothing about this!

Yes, mathematicians are. And yes, most of my colleagues did not even know the opt-out was an option.

Question is what does that button do.

I bet a lot of lawyers are salivating at this question too.


If I am reading your question correctly you are asking about chat interface Vs Codex/Claude code? If so, in my experience Codex/Claude code use is widespread for mathematicians who are seriously using these tools.

Using the space of an entire wafer for one chip would result in extremely low manufacturing yields. Even with state of the art silicon cleanrooms, there will still be defects in parts of the output.

With CPUs and GPUs, chip makers can disable faulty cores and bin them as lower SKUs to get some yield out of it. But if you're using an entire wafer to embed weights, and a speck of dust causes a printing defect that makes the weights wrong, the entire wafer is worthless.


What's the difference between disabling faulty cores and disabling the parts of the wafer that have defects?


I'm not an expert, but I think those are the same thing. But for an LLM etched onto a whole wafer, it doesn't make sense to disable part of it since that would remove some weights entirely.


Do failed wafers have to go in the trash, or can you recycle them?


You can grind some of the raw silicon out of a finished wafer but I don't think it'd be suitable to use in another batch of product. So instead of having the weights on the wafer like OOP was suggesting, hardware inference has been trending toward having a wafer of lots and lots of smaller cores with fast SRAM.


Is that defect easy to detect?


The CloudFlare blog discusses that idea when they talk about having an "agent process" to hold cryptographic material, but they list drawbacks like having to develop two processes, implement a well-defined interface, and enforce ACLs. I'm not convinced that "developing two processes" is a reason not to do it, since the kernel is effectively just the second process now, but everything else makes sense.

It's unfortunate though since this is one thing I think Windows does decently well. The Windows crypto and TLS APIs do use a key isolation process by default (LSASS) and have a stable interface for other processes to use it [0]. I imagine systemd could implement something similar, but I also know that there are very strong opinions about adding more surface area to systemd.

[0] https://blackhat.com/docs/us-16/materials/us-16-Kambic-Cunni...


TBH LSASS is privileged enough to be a good target for exploits.


Interestingly, Microsoft has been trying to get ahead of this for a couple of years now with their National Partner Clouds program [0], which they describe as:

> designed for scenarios where full ownership and operational independence from Microsoft is required

In France's case, Capgemini and Orange have a joint venture to operate datacenters that Microsoft runs Azure and Office on top of [1]. Moving away from Windows and Teams would still reduce their dependence on Microsoft substantially. But if the core goal is to reduce dependence on non-European suppliers, I would be wary of the French government buying services from "Bleu" when it's mainly Microsoft and a couple of consultancies in a trenchcoat.

[0] https://learn.microsoft.com/en-us/azure/azure-sovereign-clou...

[1] https://www.capgemini.com/news/press-releases/capgemini-and-...


I've been trying uBO Lite myself for a few months, and anyone who uses YouTube will absolutely notice that it's worse at blocking. Lite tends to delay playback at the start of a video for as long as the blocked ads would've been, making the site feel slower, and once in a while an ad will slip past the blocker anyway.


I am not so sure if that is the light version. In my (outdated) Ungoogled Chromium which still has classic uBlock, YouTube videos also have delays or do stop playing completely after a few seconds. So I have switched to the FreeTube software to watch YouTube videos. I can recommend that.


I have used Youtube and uBlock Origin lite for the past couple of months and have not noticed that. Are you using the complete filtering mode?


Just use Freetube to browse Youtube. It's a better experience in every respect.


Are you thinking of the Windows 3.00 Working Model?

https://betawiki.net/wiki/Windows_3.00_Working_Model


I'd not seen that one before, but I could have done with that version on my twin floppy 286 as it was a pain to run 3.1 on it.


Not the whole thing: a lot of the Windows org chart is still under Rajesh Jha in Experiences + Devices, or scattered around Azure with Scott Guthrie. But they've already been pushing Windows Copilot and Bing Ads and widgets, so I imagine the plan is more of the same.


Mikhail reported to Rajesh Jha and I guess now he'll report to Mustafa.

Mikhail's official title is CEO of Advertising and the entire Windows Engineering org reports to him.

I just asked some MS friends who confirmed.


It's not just the kids growing up now. I'm sure plenty of millennials who watched YouTube when AudioSwap was a thing will recognize "Dreamscape" by 009 Sound System:

https://www.youtube.com/watch?v=TKfS5zVfGBc


> "Dreamscape" by 009 Sound System

Say no more. I'll go get Unregistered Hypercam 2.


I think the IPv4 "evil" bit [0] already does this :)

[0] https://datatracker.ietf.org/doc/html/rfc3514


Whether it's proper depends on who you ask but you can use the singular they.

https://en.m.wikipedia.org/wiki/Singular_they


This is interesting.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: