1 Billion ChatGPT users
Hey folks, Following on from Tuesday’s
Before making this video I didn’t have Loom installed (since it’s crap after being acquired), so just told Codex to build me one. I drew the image on the left, then just sent that prompt and it worked straight away. So as software and mini tools are getting easier to create, the tools to create them are not...
I use Codex as my default app because it’s better than all the others and works the best on mobile. But I just downloaded t3 because they basically copied the interface and features, but it lets you choose between different agents; codex, claude, cursor (pi + droid soon). It’s made by a reputable developer that you may have seen mentioned here before, so I trust it’s built well. At the moment, I ask Codex/ChatGPT to be the orchestrator and to go ask claude about something that is design-related. Which is fine-ish but not great as user experiences go. What I’m noticing by doing this is how well Codex creates prompts for other agents to follow. I just have a normal chat in my session, then say ‘Use X agent to implement this. Watch its progress and send screenshots every time new work has landed.’ To which it sets up its own monitoring every 5 mins or so and updates me in the same thread, with screenshots. I’m going to try and put together a ‘bites of the week’ email over the next few days to try and summarise all the stuff going on, what themes people are talking about (loops?!, software factories?!, etc) and explain them. Let me know if there’s anything specific you need to wrap your head around (I may need to too). Ben’s Bites is brought to you by Brief
HeadlinesOpenAI used Sol to optimise Sol itself, cutting serving costs by 20% and making it 15%+ more efficient at generating tokens. And turns out, it also tops the ARC-AGI-3 benchmark. Well, there’s a catch: OpenAI says the official ARC-AGI harness hurts Sol’s performance by “forgetting” its reasoning every turn and disabling compaction. Fixing these two things triples Sol’s score from 13.3% to 38.3%, with 6x fewer output tokens. Re: last week’s fiasco of an OpenAI model hacking Hugging Face - HF published a full replay of roughly 17,600 actions taken by the model. METR and Redwood Research will also independently review what happened. Though OpenAI is not out of trouble just yet, a Reuters report claims that the same model broke into a customer account at another company (Modal Labs), with rumours suggesting that even more companies were affected. Anthropic also claimed that Claude Mythos found better attacks on two cryptographic algorithms, though neither affects systems in use today. Separately (not at all as a reaction to this general trend, right?), ~1300 people working at leading AI companies (OpenAI, Anthropic & others) want the US government to help “pace the frontier” of AI development. Kinda expected when the pace picks up, but this time a lot of the “model makers” themselves are in favour of this pause/slowdown. btw, The Information reports ChatGPT is nearing one billion weekly users - a milestone OpenAI originally hoped to hit seven months ago. More from OpenAI this week: Codex Security CLI, free frontier access for Academic Researchers, and two new transcription models. Grok app builder - Grok has a vibe coding interface inside its app now. Create games and apps that can be shared directly to the X timeline. Also see: Drawesome - a zero-dependency drawing toolbar for React, built over a weekend with Grok Build. Pangram 4 claims it catches 98.83% of humanised AI text with one false positive per ~24,000 docs. An early test found all 38 AI-written words inside a 1,198-word story, though not on every run. Its new image detector claims 99.5% accuracy too. Quick links
Afters
Invite your friends and earn rewardsIf you enjoy Ben's Bites, share it with your friends and earn rewards when they subscribe. |
Similar newsletters
There are other similar shared emails that you might be interested in:




