I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines.
I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries
I feel like there are so many cloistered people at these companies that they are left scratching their heads about what normies even want. Like, they literally can't fathom basic stuff that isn't just highly consumer-oriented. I dunno, like applying for government services, paying your gas/elec bill without being confused af, keeping the dr up to date with your dad's illness, or how to get your newborn to sleep at 2am.
I feel like we read different blog posts. The announcement shows Astra filling out a tax form, looking for a kindergarten and apartment hunting. These are ordinary everyday tasks for normal people. It also shows lots of non-consumer stuff like financial modeling, genetic sequencing analysis and coding.
I will happily shout AGI from the rooftops the day I can turn on voice mode in ChatGPT and have the model calm down my toddler for a tantrum, or keep him from opening all the bananas instead of eating one.
People in this thread arguing about AGI relating to Einstein problems and physics. Yeah no. When AI can handle toddlers then we're getting somewhere.
It's actually a really interesting question to reckon with - can an AI help raise a child and is that good or bad? I suspect there are ways you can use the current frontier that elevate the child's experience, especially when it comes to education and creative play. On the flipside using one of these models to automate storytelling to your 3 year old is probably a bad idea... The tricky bit is where to draw the line? Is using an LLM to collaboratively build a story alongside the parent and the child bad? I don't know, I suspect not, but that is not a simple question to answer.
It all comes down to whether you see time and attention spent with your kids as a nuisance or the best time you can ever have in your lifetime (I am team B).
And as every child psychologist says, "attention is all you need" :).
Well there is an upper bound of how much time you have available to spend with your child and how you spend that time. You could use an LLM when you can't be there (I suspect this is a bad idea), or you can use it collaboratively with your child (I suspect this is a good idea). I see it as similar to using a phone; your child on wikipedia is a different thing to your child on tiktok.
We have neither AGI nor _things_ I'm afraid. I wouldn't even know what to do with a semi-autonomous automation tool for my daily tasks, I don't really have anything that I'd like to automate.
(like, how often do you look for a doctor or day care? The demos automate occasional and one-off tasks. That said, it automating tasks like circuitboard design and CAD modeling is impressive and will / may have a much higher impact on those respective industries (if it works), just like LLM assisted coding did)
Many VCs also dislike these examples, I believe. I'm doubtful this is what they're being pitched.
As for public releases: I wonder if it's because these examples are easy to relate to. Many websites are just a long tail of industry or use-case specific stuff. What's valuable to me probably means nothing to you. This is unlikely to resonate with people-wit-large (and LLMs are marketed broadly) or requires the reader to think (and marketing that requires thinking is bad these days).
Second, it's arguably a good litmus test. If it still can't do the worn out examples of plane tickets and shopping, which would be a good assumption since we've been demo'd these use-cases for 2 years at this point, then ...
They’re ridiculously facile use cases for what agentic AI can already do.
I don’t trust an agent with full executive power - yet - but I am essentially using GPT as my PA. Right now I’ve got it managing a construction project with recalcitrant contractors, an international move and visa tied to a property purchase, a short term holiday rental business, and basically just popping up in my life going “the situation is this, you need to do X/I suggest you send Y to Z, the email is prepped in your drafts”. I am of course also using it for software development, and have had it resolve every digital chore in my home life.
I guess it doesn’t make for a quick elevator pitch, but I’m finding it has reduced my cognitive overhead on a whole raft of fuckery that would otherwise have me shouting at inanimate objects.
Anyway, today I have to take the cat to the vet, email a bank, go and discuss ceiling systems, and review a proposed shopping list for a maintenance visit to a holiday rental. Can’t wait for the bleeding robots to get here.
Also travel plans. For some reason these marketing guys love to post pictures of places to visit on a board and share them with their friends list. Then a friend makes a comment about a picture on the board like "Cool, I can't wait to be there <emoji>". I've always assumed that's a cultural thing.
The enthusiasm from investors is based on the premise of replacing all keyboard jobs (programmer, lawyer, accountant, etc.), but it is not a good look to talk about this to the general public (who might be programmers, lawyers, accountants, etc.).
AI heaven is using a computer to do your job. AI hell is your boss using a computer to do your job.
There's a lot of truth to this: Execs want their own use case covered. In addition, everyone wants to be better than the competition at basic use cases, because that's what people compare first - even though we all know they are not a good reflection of reality.
I'm currently working on exactly a project like this.
And all of these things can easily go wrong with this halfbaked "AI" technology - getting thow wrong tickets you can't cancel, loosing all your money due to bot buying random titanium cubes & deleting your emails!
All those stupid Bell Labs researchers not inventing Uber or Tinder. How come they didn't just build the obviously popular and profitable businesses that became possible once they invented the internet?
Compare that to Xerox, for example. Pretty early on, they had visions to replacing / enhancing largely "common" use cases (basically, everything that
required printing a letter or sending an mail) [1]
I don't know if, at the time, people where doubting that it would be useful. (Practical ? Affordable ? Other legitimate questions.)
But here, the contrast between the promises ("it will cure cancer and solve climate change") and the demo ("it can cost you money on stuff you never asked to buy") is a bit telling.
But I agree that you can't expect the enablers to think of all the uses cases and applications.
Still, just so I know: what ARC-xyz score means the LLM has cured climate poney cancer, exactly ?
That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?
Corporate travel is an example. In many organisations, you tell someone in the travel department "I need to be in Tokyo for this conference from Tuesday to Sunday, and charge it to this cost code", and they figure out flights, accommodation, etc for you, with minimal input from you.
Unless this someone in the travel department knows you and has booked many flights for you before, there are so many unstated assumptions and preferences and hidden utility functions that make that someone unable to find the best solution for you.
My employer does have corporate travel concierge who can actually do this with minimal input. But people do not want that. They think about aspects like, if I have a layover at XYZ, would the vicinity of XYZ be a good place for a few hours of sightseeing, given my own travel preferences and previous sightseeing destinations; and would I still want it if it is an overnight layover at XYZ. Or things like, if I fly via this route, I could squeeze another paid day with zero work expectations while I could rest on the plane / in lounges. Or things like, this particular airline has this particular type of aircraft with far better seats than that particular airline, so I would be happier even if the flight is longer.
Needless to say, all colleagues that I know of (myself included) prefer to book corporate travel themselves, even if from the company's perspective, spending a few hours to do it is much more expensive than having the concierge spend a few hours to do it.
I agree, but often corporate policy overrides most people's preferences, which is why such processes exist - to ensure that travel arrangements match corporate policy rather than individual preferences.
In my experience corporate policy is primarily built to control cost. It will not allow booking expensive flights when cheaper flights are available. But when prices are similar, it doesn’t really dictate much.
And in many cases having a corporate policy means just adding a few more constraints: it makes the problem more like a logical puzzle and this can nerd snipe people.
Ha, but I always hated when my company tried to get me to do that and ended up spending more time on the phone with the "concierge" than it would take me to do it myself on expedia. And I think most people feel similarly.
When I tell a bot to find the best value per volume for a reasonable quantity of unscented Dawn dish soap [so I can buy that], then: It often makes a complete mess of this seemingly-simple operation.
(And yeah, that is an actual thing that I've tried to accomplish with voice commands while standing in my kitchen and doing some dishes. It seems very simple, and it did not go well.
Maybe when we get the basics figured out we can start worrying about how inept it is at doing vacation planning.
It seems that this kind of thing isn't sorted at all, and that this is a very real problem for those who are in the bot business: These missed opportunities leave money on the table.)
I think we can safely assume that AGI is still a good ways out then.
These AI LLM products really do well on the illusion of seeming intelligent to regular folks, but I'm not fully convinced they can replace human brains... yet. ;)
I mean: That sample was only demonstrative of shopping for dish soap. :)
I often get some seemingly-good results in the non-shopping research department, where I'm exploring science and physics; things that don't generally change rapidly.
But unlike dish soap, I can't buy the science results that I want[].
[]: er. well, ackshually... let's just not talk about that concept right now.
I wouldn't even use a bot for that, there's price comparison websites and of course I know where to find the cheapest stuff already.
Which makes me think: would a generation raised on AI and automating these things even consider looking for e.g. deals on household consumables? This is spiraling doomerism at this point, but, do people who outsource the hard thinking to LLMs get fresh ideas or inspiration still?
I mean: I've been using price comparison websites since we've had a web with enough prices to compare them, and they always find some way to fail. And I don't possess the hubris required to think that I always know where to buy the cheapest of anything.
And the only reason I'm thinking about buying dish soap is because I'm washing dishes. Without some kind of other influence, I won't be thinking about it again until the next time I'm washing dishes. But meanwhile, there I am -- washing dishes, and thinking about buying dish soap. I'm stuck doing this job until it is finished and my hands are wet; I might as well engage the bot.
> Which makes me think: would a generation raised on AI and automating these things even consider looking for e.g. deals on household consumables?
Sure. If finding a deal on a household consumable is what they're behooved to do, then why would they not? People broadly do adapt to the world they grow up in, and to use the tools available to them.
We don't use phone books or video rental places anymore. We're still getting by.
And people will misuse the machine as well, but that's also not new.
When the engine light comes on in ~any car from the last 30 years and the dude at Autozone plugs a widget in and declares "It says here that it's an O2 sensor!" and people blindly treat that as a firm diagnosis, then they sometimes find disappointment: They install a brand new O2 sensor and things don't improve at all. That won't change.
(In reality, that result just means "The measurement reported by the O2 sensor is irrational." That could be because the sensor is bad, but it could also be because any other part of the fuel+air system -- from front to back -- is being funky instead.)
> do people who outsource the hard thinking to LLMs get fresh ideas or inspiration still?
As much as they ever did, I suppose. Not everybody is naturally inquisitive, nor gets to be above average.
I wanted some banana plugs for my speakers, as well as I needed some longer speaker wire. I kinda wanted it for the weekend so I could watch movies with some friends. I didn't want to order online as that wouldn't arrive in time. Three stores around me had some, but they were quite far apart and I didn't want to drive to all of them, so I needed to figure out which store had ones that were actually decent and not too crappy.
LLMs completely and utterly failed at this.
The Mensa test results are impressive but are in no way shape or form relevant to real world scenarios.
I wanna see benchmarks on successfully completing benefits applications, on finding the best option to purchase X from, on correctly identifying a piece of furniture and its condition and setting a correct price on it and selling it on craigslist/FB market place and so forth.
I've been looking for usable, insulated, stackable banana plugs that work with 12 AWG wire for over a decade.
For a lot of my audio things, I can get by with MDPs from Pamona Eelectronics. They're insulated-enough, stackable, dual-banana plugs on 3/4" centers and where they work, they work very well. (I don't recall if 12AWG is within spec or not, but it can be made to fit.)
But they do not work on audio widgets that do not have 3/4" spacing, like my aging Lexicon receiver and its great quantity of binding posts for speakers.
I don't care if they're glitzy or fabulous-looking. I just want functional, insulated, stackable single banana plugs.
(The common marketplace is full of elaborate gold-plated non-insulated things. I guess they're pretty, but they're a hazard and nobody will see them again soon at all if everything goes well. Their function-follows-form beauty is only detrimental.)
Funny, I had a similar task for an LLM recently but it succeeded, so I'll describe it:
I asked the chatbot (librechat, tavily-mcp connected, kimi-k3) if anybody tests dish soaps rigorously. It said yes. I asked it to find the consumer reports from Germany. It did so, full access is paywalled but it did see the top positions. The I asked it what term to put in evay.de if I want to order. It gave me a string that found the auctions for me. I also asked it what the actual difference is and it described what info the EU law forces onto the label and told me what to look for in the ingredient lists.
The process is still involved. I wish I could just do it all by saying "find me another dishsoap, this one sucks". But then would everybody saying this phrase except the same approach to the issue that I took?
I was about to argue against this with something like "soon everyone will just follow the average Claude/GPT generated trip to every destination"...
But I realized that this is mostly a strict upgrade to how many many people are planning trips these days. It will probably do a pretty good job of accounting for your preferences, probably better than most/any travel agent unless they're very very local.
Yeah but the contention is here automating the bookings. Did Claude book the flights, hotels, tickets and attractions for you or would you let it without oversight?
Yes it booked everything, I just had to click the final pay button. I did run my itinerary against a fresh session like: 'There are errors in my travel plans, go find them'. It found one issue where I was scheduled to be at the wrong train station.
I think it depends on what you do for work. I'm not going to ask an agent to book my flight for my vacation to French Polynesia. I want to pick my seat and potentially find a deal making an upgrade worth it, choose an airline, etc.
But my routine business trips in the CONUS with strictly defined booking options... let me just email an agent "Get there by meeting on day A, leave after meeting day B" and have it sort it all out without the drudgery of the corporate travel portal. YES PLEASE!
By that point, you don't need an AI that boils the oceans and hopefully doesn't confabulate or misinterpret your words, all you need is a better corporate travel portal, which is a legitimate workplace productivity discussion to have, and a solved problem with traditional methods. Ours is effectively close to the workflow you describe: picking dates and time brackets, destination, fine tuning flights and hotel options, sending for validation, less than 10 clicks through information-dense and effective/predictable screens which I wouldn't want to trade for a chatbot and it's usually over the top words salad.
We have had idea of expert systems for decades at this point. And spend however much on software development. Some how we have failed to reach this in way too many places. Even simplest things like cancelling some service might be broken...
And now we want to add some random factor into middle of it all...
And the reason why there are 10 clicks is that that is the number of decisions that is required to get people to hand over their money.
There will have been an A-B test with 9 clicks -> less sales, less revenue.
There will have been an A-B test with 11 clicks -> less sales, less revenue.
The people making these sites are not dumb. They are also sitting on thin margins nowadays, and mostly from the airlines and hotels rather than from the users.
That doesn't apply to a corporate travel portal. If you're on the portal, you're booking a business trip. You're not going to not book a trip because it takes 1 click too many.
That's approximated by the portals - the people building these go and look at Expedia for their flow and design ideas... so the information permeates across.
For me, it is not a matter of trust but that I actually like shopping, planning a trip, deciding what restaurant to go to. Deciding what to buy when shopping is a matter of personal taste and not intelligence.
A human assistant is largely a status symbol. Most people are not really that busy. The real problem with an agentic assistant is if everyone can have one then it no longer acts as a status symbol.
And I could bet money on that in short time after someone builds that sort of system the next step is to make it worse. Push worse and more expensive options to user. Or at least those from highest bidder... Anyone involved just can't keep themselves honest so it is doomed to be exploitative.
Plan a holiday, definitely. There are many esoteric things one has to research to properly plan a holiday that I’d rather just not. Some things require reservations months in advance and I’d rather an AI just figure all that out for me ahead of time.
You can't even get many people to buy things online at all and if you can it's less profitable than retail, because you need to spend a lot of money to convince people, advertise to be seen, and account for returns. I think this is also due to the factors you mention.
One quick example: In fashion, Inditex and Shein have about the same revenue (€39.9bn and $41.8bn in 2025), but Inditex is more than three times as profitable. I don't see how there is a demand for agentic commerce that would remove even more control from the customer when shopping. Part of why we shop is for the experience. For B2B producurement platforms like Alibaba I can see the appeal though.
I recently needed to buy some hardware for a piece of furniture.
Ran Codex, it found it for 18% less than what I found in the top Google results. It did it by finding smaller shops, applying a discount code, subscribing to a newsletter for a better code after approval, and took into account the shipping (by placing it in the cart and going to checkout) all to get me the best price.
I’m guessing without it I would have spent much more time on it and paid the original price I saw.
If you use AI agents well, they can easily save you more money than they cost, and saving money is something most people are pretty excited about.
Hi Omer,
I think that is a use case where I'm sure an AI agent does well
right now, but I have 2 follow-up thoughts based on it and what
it is means that this is one of the best use case for AI online
commerce so far.
My first thought was that this is a quite limited application as
most of shopping (online & retail) is not spent comparing prices
and is instead spent on finding what product to buy in the first
place. I believe this "shopping fun" and exploration of multiple
options is not going away any time soon as those are some of the
most attractive parts of the whole experience.
I also think that the price comparison websites that exist, e.g.
geizhals.de here in Germany (includes most e-commerce shops, and
shipping information), are pretty good already and AI only being
a better version of that would be quite a sad turn of events and
with ads (and commission systems?) potentially coming to ChatGPT
in the near future the incentives to find the best price are rly
misaligned and I have doubts that people will trust results from
ChatGPT. We already saw this as Google's sponsored links evolved
to show whoever spent the most for the placement. Paid providers
without ads and commissions might fare better though.
I mean, that's cool and all, but the numbers are really going to shift when it accidentally goes off and orders that same hardware from every vendor in your local region and the top 5 online results for comparison.
It's the same problem as all other LLM solutions (that I hope OpenAI is working on!) it's non-deterministic, and there's no way for the user (or model provider) to know what the distribution of possible outcomes is. This just gets compounded when multi-call harnesses come onto play.
I‘m using mostly the browser automations for things like that now. Same for research - ChatGPT is banned from reading many pages, but Codex can read anything I can.
But isn’t it funny that Cloudflare is blocking AI on their pages, but on the other hand is researching and marketing things like „you can put a browser in a CF worker“
Their browser workers won’t get blocked. Same with all big vendors, keep out small competitors enjoy access yourself and sell it to a select few partners.
I expect a small site that undercuts the top Google results by 18% with a sign-up discount probably isn't profitable on those orders, so blocking agents would save them money - it's not like someone using agents in that way is going to have any loyalty to shopping from that site in the future.
Yeah, I mean if that's the case, then said small sites are probably headed for extinction, and agents won't have anything to work with other than whatever the top Google result is anyway.
And yours and other's agents will probably remain an insignificant and invisible customer base anyway.
My crystal ball is as good as anyone's, but if "agentic shopping" ever becomes mainstream, you can be sure that the vast majority will ask their phone (i.e. Google, i.e. Google Shopping) what the best price is anyways.
Sure, Google could be their agent, or ChatGPT, or Claude, or whatever local model the person is running. The big guys might all have the results cached so they don't have to rerun the crawl. Whatever their choice of agent, it seems pretty clear that almost everyone's going to use them, they're way too useful not to.
True, but they're still friction to be reduced here.
What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.
I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
I work for a larger german retail chain and agentic shopping is already on the "near future vision".
No one thinks this will be used but somehow shareholders love it.
I think that one thing that "thinkers", academics and other smart people who LIKE to think/select/choose forget is that the majority of the population is not wired this way.
Most people go about their day absolutely minimizing the amount of mental energy they have to spend. They have priorities like kids, work, family, groceries, etc. If anything here can be automated its a fat win in their life. I have friends that are loving the features where meals are planned for them, food is home delivered, Uber is auto ordered, trip plans are made, etc. They don't mind paying more just to reduce mental load. Its a huhe market.
Maybe we are not the real audience. I wonder if they are trying to convince the advertising industry/investors that in the future it won't be Google search that stands between the consumer and the product, but rather their LLM.
The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint.
I wouldn't want it to pick food for me from a place I've never been, though to be honest with enough order history it could probably do a decent job at it.
Agreed. There's not many things I don't want AI to help with, but buying stuff autonomously is high up on the list of things I don't want. Brockman's latest interview was something like: "AGI would be able to say oh this band is playing, I bought the tickets for you and arranged your flights - I hope you don't mind" (paraphrasing here). I definitely don't want AGI running my life like that so I can be a mindless consumer. I'm sure the advertising/marketing companies would love it though, so they can make closed-room deals with AI providers to shill you garbage you don't need. Just another reason why open-weight models need to keep up.
The core or the problem is also what you are describing was bought up by Yuval Harari in his interview with the economist.
To paraphrase some of his lines: We make decisions using our emotions and our thoughts. What makes us different from the AI is that we can be afraid.
To portray a guy as ordering "beef bulgogi", in the same breath as "email this rocket design marketing", while it _might_ seem appealing and resolute, though oddly fast paced, seems pretty ignorant of the human _quality_ that make most practical decisions messy.
Maybe because the people making them are workaholic types who really don't care? I've certainly been in situations where I didn't really care what shows up for a meal. Someone was tasked with getting food and we leave it up to them. Sure, I can imagine such a meal being bad and it has been a few times in my life but 24 of 25 times, maybe more, it's fine. Further, the AI knows your preferences.
(This holds generally) maybe because the people most likely to be convinced to buy this particular thing are the people who buy a lot of things. The people who only buy things they need, they’re a harder demographic to convince via ads, and they tend to pay less since they rarely need extras. Every ad caters to whales.
That's not far off the ads for smartphones with things that I never see people doing in real life.
Currently there's a Google Pixel ad where a grandma takes a photo of a board and Gemini automatically fills her calendar with all the events. Yeah sure.
I really wonder what the demo people are thinking as there are many great uses of LLMs in day-to-day lives but all we get is these 3 rehashed use cases. No wonder normal people think LLMs still can't do anything.
"Somebody is wrong on the internet" will get you more, and more reliable, results on what people actually want from AI. You are now participating in a very high-value survey.
The point of advertising is to normalize the behavior that makes you money. Take cars for example, millions of people have been killed by cars and they are objectively poor for public transportation yet through a few good advertising campaigns and lobbying efforts American society redesigned every city cars.
About half the oil the US uses (for the last 50 years) goes to cars so you can put every middle east boondoggle onto the bill too. Don't think that just because it's the stupidest idea you've ever heard that people won't rejigger society to make is kinda work for 50 years.
I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate.
"Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip!
I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming up with such demos.
I hooked up Google Maps to ChatGPT and that was definitely helpful for this; without them it would seemingly just guestimate driving distances, but with the MCP it was able to determine exact driving/walking times and also render maps with waypoints for activities around the lodging.
> why do so many of these demos include people buying things autonomously?
> Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
On one hand I'm with you. On the other I thought making ecommerce purchases on your phone is absolutely idiotic idea that will never catch on.
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)