Or AMD’s putting sram on a separate chip. I’m surprised they didn’t go there yet except for extra L3 instead of all of it. At some point the costs will shift the decisions.
I know a freelance translator and they have more business than ever, AI being "mostly" right is not good enough in many businesses. You need oversight and most importantly - accountability in the form of a human that can put their stamp of approval.
Sending email is already cheap, not free. Just so happens the cost is subsidized/hidden for most regular people and its so cheap that they give it away for free. For those of us that run transactional email, the cost is ~ $0.0001 per email for AWS SES, already at the fraction of a fraction of a cent you suggest.
Yes, but I want the rate to be variable depending on volume. I want it to be $0.0001 for a simple personal email. But if you send 10,000 emails today, I want it to cost you $100 (or more).
Many schemes to do this. Things like:
The first 10 emails is $0.0001 each.
The next 10 will be $0.001 each.
And so on. You can always fiddle with the thresholds and costs.
Anyone have any idea what the architecture/vendors they are using for inference/compute?
Getting the compute to run inference for multi-trillion parameter models at any sort of scale and performance is daunting. There are a handful of vendors that have systems that can do this (~ Nvidia NVl-72 class) that pretty much only the frontier labs and hyperscalers effectively have access to.
That statement means nothing. You could say the exact same thing about Rails and have an equally defensible position. What about its architecture makes it better?
reply