Leveraging AI
Dive into the world of artificial intelligence with 'Leveraging AI,' a podcast tailored for forward-thinking business professionals. Each episode brings insightful discussions on how AI can ethically transform business practices, offering practical solutions to day-to-day business challenges.
Join our host Isar Meitis (4 time CEO), and expert guests as they turn AI's complexities into actionable insights, and explore its ethical implications in the business world. Whether you are an AI novice or a seasoned professional, 'Leveraging AI' equips you with the knowledge and tools to harness AI's power responsibly and effectively. Tune in weekly for inspiring conversations and real-world applications. Subscribe now and unlock the potential of AI in your business.
Leveraging AI
298 | $500M token charge?! 🤯 The AI race is changing, Claude Opus 4.8, a robotic house cleaning service and more AI news for the week of June 5, 2026
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Secure your spot for the MULTI-AGENT ORCHESTRATION AI COURSE: https://multiplai.ai/multi-agent-orchestration-course/
Are AI jobs disappearing faster than they're being created—or are we asking the wrong question?
For months, AI leaders warned of massive job losses. Now, some of the same voices are changing their tune. But while Sam Altman and Dario Amodei are sounding more optimistic, tech layoffs continue to climb, AI anxiety is spreading across the workforce, and business leaders are facing a difficult question: what's actually happening beneath the headlines?
In this week's AI News episode, Isar Matis breaks down the conflicting signals shaping the future of work and explains why the next few years may be far more turbulent than many expect. He also explores a major shift in the AI race: why model intelligence is no longer the primary battleground, and why speed, cost, accessibility, and business value are becoming the metrics that matter most.
If you're a business leader trying to understand where AI is headed—and what it means for your workforce, strategy, and competitive advantage—this episode provides a practical framework for separating hype from reality.
In this session, you'll discover:
- Why Sam Altman says he may have been wrong about the pace of AI-driven job disruption.
- The Jevons Paradox and how increased AI productivity could create new demand.
- The reality behind recent tech layoffs and whether AI is truly responsible.
- Why employee anxiety around AI may matter more than the statistics themselves.
- The growing gap between jobs being eliminated and new AI-related roles being created.
- Why reskilling workers may be harder than most organizations expect.
- How the AI race is shifting from model quality to business value.
- Why open-source models are rapidly closing the gap with frontier AI systems.
- The hidden cost explosion companies are experiencing with AI tokens and agents.
- What Microsoft's latest Build announcements reveal about the future of enterprise AI.
- Why robots cleaning homes may be an early sign of AI's impact on blue-collar work.
About Leveraging AI
- The Ultimate AI Course for Business People: https://multiplai.ai/ai-course/
- YouTube Full Episodes: https://www.youtube.com/@Multiplai_AI/
- Connect with Isar Meitis: https://www.linkedin.com/in/isarmeitis/
- Join our Live Sessions, AI Hangouts and newsletter: https://services.multiplai.ai/events
If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!
Hello and welcome to a weekend news episode of the Leveraging AI podcast, the podcast that shares practical ethical ways to leverage AI to improve efficiency, grow your business and advance your career. This is Isar Meitis, your host, and we have a lot of really interesting things to talk about this week. We're going to start with a deep dive into what is really the impact of AI on the job market. There are contradicting forces right now, and we'll try to put everything in order and provide all the different aspects and views, and I will give you my opinion at the end of it as well. We're also going to dive into what is changing and what really matters right now in the AI race. There has been some major shifts and some new focuses from all different directions that are going to impact how this race moves forward, so that's gonna be our second deep dive. We also have a lot of announcements from Microsoft from everything they announced on their Build event this week, and we have a new model from Anthropic, new capabilities from OpenAI, And we have robots cleaning houses in San Francisco. But before we get started, I owe you an apology for not releasing a news episode last weekend. I was traveling internationally, and this week I Had five different workshops to four different groups of people, which took a lot of time to prepare for and finalize all the details, and combine that with the fact I was finally meeting my family overseas after a very long time. I just did not have the bandwidth, to also record an episode. So again, I apologize for that, but we have a lot to cover this week, so let's get started. As I mentioned, the first topic today, we are going to dive into what is really happening in the job market, or more importantly, what we think is going to happen in the job market when it comes to AI. So both of the main leaders of the AI race right now, Dario Amodei from Anthropic and Sam Altman from OpenAI, have shared previously multiple times that they see the future of current work as doom and gloom. But this past week, both of them sounded different voices. So Sam Altman, in an interview with the Commonwealth Bank of Australia's CEO Matt Comyn, has said that he believes he was pretty wrong about AI economics, And when he was asked about the impact on jobs, he said the following: I'm delighted to be wrong about this. I thought there would have been more impact on entry-level white-collar jobs being eliminated by now than has actually happened." We also have seen Dario Amodei, who previously said that AI could eliminate 50% of white-collar jobs, now backpedaling from that and saying something a little different. He said, "If you automate 90% of the job, then everyone does the 10% of the job, and the 10% kind of expands to be 100% of what people do and kind of 10Xs their productivity." Or if you wanna put that in clearer and easier to understand terms, he's talking about the Jevons paradox, right? Which says that when a cost of something, of a service declines, then it doesn't reduce the need for that, but actually expands the demand for that because now the economics of every unit is significantly cheaper, which drives more demand. This has been a known phenomena for decades. We also had additional voices pointing in the same direction. David Solomon from Goldman Sachs and Torsten Slok from Apollo, went back and looked at the history. Solomon invoked the 1900s electrification and 1990s digital revolution as precedents to saying that as technology grows, It actually grows the percentage of employment rather than shrinking it, noting that the US employment grew 145% since 1962 to today, despite all the changes in technology that presumably have enabled more automation and reduced efforts of humans in the work that they're doing. And Slok said, "Lower cost per interaction does not mean fewer interactions. It means more customers served, more channels opened, and more markets worth reaching." And obviously, there's truth in that as well That being said, going back to specifically Sam Altman and Dario Amodei, they're both on the verge of an IPO this week. Both might be two of the largest or maybe the two largest IPOs in history, and both need a positive public opinion and tailwind in order to make their IPOs as successful as possible, especially when these two IPOs are happening in the same summer with the IPO of SpaceX. So three gigantic IPOs all in the same few months are going to really stretch the liquidity of the markets. On the positive public opinion and the payment from an investment perspective. So there's definitely some hidden agendas, or maybe not so hidden, but clear agendas why they're saying what they're saying now versus what they said a few months ago. But there are definitely more and more voices that are saying that the doom and gloom future that they thought is not, which is at least what Sam is saying. He didn't say it's not going to happen. He said he's surprised it hasn't happened yet. However, when we look at real numbers, we have seen 2026 tech layoffs topping 115,000 jobs just through the end of May, which is almost the full number of 2025 with 124,000 tech jobs lost. Now, Challenger, Gray, and Christmas have done a research and they found or they attribute 54,836, that's a very specific number for research. 54,000 jobs, almost 55,000 jobs are going to be lost in 2025 alone that are related directly to AI. That is compared to a total of 72,000 since 2023. So 2023, 2024, and 2025, the total number was just short of 72,000, and they're expecting 55,000 just this year. And they're saying, based on their research, that about 16% of all job cuts were attributed to AI. Whether that is true attribution or not is very hard to know, but that's what they're trying to find out in their research. If we just look at the last couple of months, we have Meta and Microsoft combined letting go of over 20,000 people. And if you combine all the numbers from some of the bigger tech companies, we are nearing 90,000 tech workers that have been laid off just since April, so just in the last two months. With some leaders specifically attributing it to no need for more people because of AI, a very prominent one was Marc Benioff, the CEO of Salesforce, who said, "I've reduced it from 9,000 heads to about 5,000 because I need less heads," and he was referring specifically to customer service agents in his company. I wanna pause for a minute and talk to you about our courses that are coming up. This episode is brought to you by the Multi-Agent Orchestration Course. This is a course I started teaching recently. We've sold out all the courses between May and the end of July, so the next course we are selling for the next cohort starts at the beginning of August. It is a four-week course that will take you from knowing nothing about what skills and agents are, to learning how to do it at a production level grade, including everything you need to know on how to build them, how to connect them, how to orchestrate them, how to build the infrastructure for them so you can develop them and deploy them effectively. I'm also teaching these courses as workshops for companies, private workshops. So if you're in a leadership position in a company and you don't wanna just take the course that everybody else is taking, and you want it custom-tailored to your company's needs, please reach out to me. if you want to take the courses, there's a link for that in the show notes. Just click on the link and you can sign up. If you are doing this before the middle of June, you're going to get $100 off. So when we sold the first courses, it was $500, then $200, now it's $100, and that is going away at the middle of June, and there's not going to be any more discounts. And most likely, we're going to push the rate higher because most of the other people who are selling these kind of courses are selling them for significantly higher prices than we are. So if you're wanting to learn and to push yourself into the agentic era, miss this opportunity. This course will sell out just like all the ones that we sold between May and July. Come and join us. Enjoy the $100 off by the mid of this month, And I would love to see you on the August cohort. And now back to the news. Now, there are a lot of people in the markets, including some leading economists and investors known, such as Marc Andreessen from Andreessen Horowitz, who are saying this is completely AI washing, meaning it's CEOs who are letting go of extra labor that they had too many of for the last few years, and they're just cutting people off And using AI as an excuse instead of saying the real truth, that they overhired and they had too much fat in their organizations. There are also other cases like Wix this week Wix the company that helps you build websites and host your websites very easily without having any technical skills. They have just announced a thousand job cuts, which is another 20% of their workforce, while its revenue is actually growing. But the real story is that their stock has plummeted since mid last year and dropped 50% since January with a single day of 27% drop just in the last couple of weeks when they announced their results, which missed their targets by a very, very big spread. And the reason for that is what you can do with Wix, you can now do with any AI in almost the same level of ease, and people understand that there's maybe no future in that business model. So if you are the CEO of Wix and you can say, "Our business is going to be eliminated or dramatically cut by the capabilities of AI," that's not a good thing to say as a CEO of a large, successful international company. Saying that you're building efficiencies by using AI internally sounds a lot better, like you are improving the company versus trying to fight the inevitable. And so that's another example of a lot of cuts that are being attributed to AI, but there are a lot of other stories behind the scenes. We also have Cloudflare, who announced a 1,100 employees going to be let off. That again, very similar, 20% of its workforce. But in this particular case, Matthew Prince, the CEO, actually gave Wall Street Journal an interesting view on how he sees the future impact of AI on jobs. So he said the vast majority of those we laid off last week were measurers, and he divided the employees into three kind of employees: builders, sellers, and measurers. Builders are obviously the people who build the actual technology of Cloudflare, so the engineers, and sellers are the people who are the sales team, people who are selling the services, and the measurers are everything in between, the people who do not build and do not sell. A lot of them are middle management or other layers of supporting roles. And he thinks that there will be a smaller and smaller need for these kind of people over time, while there's going to be an increased need for builders and sellers because they can make more of what they are selling and sell more of what they are selling, which makes perfect sense. Again, similar to many other companies in which we've seen cuts, Cloudflare has seen a huge increase in their revenue, and they're hitting record revenue rates every single quarter, and yet they've let 20% of the people go. Another interesting data point that relates to all of this came this week from a research that was done by a company called GCheck. GCheck is a company that does background research, and they've done a survey, which I'm going to talk about how they've done it because I'm not completely excited about the way they've done the research. But what they found is still interesting. They found that 63% of US workers have lied or exaggerated their AI skills, rising to 80% for Gen Zs and 70% for men versus a lower number for women. Now, if you dive into the actual results, you see that the header they've used with a 63% also includes people who said that they are trying to sound their opinions more to sound more knowledgeable, which doesn't really mean they're lying. It just means they're trying to push aggressively to let people understand or assume that they know, uh, enough about AI. But it's still very obvious that employees are starting to care more and more about what employers think about their AI capabilities, compared to what their capabilities actually are. Another thing that they found that relates to that is that 81% of workers discourage or limit AI usage at their work because they're afraid of the implications of what that might do to their own work and to their own employment capabilities, and 53% prefer manual approaches over learning how to do things with AI. The other aspect, the flip side of that is 64% of workers have never had their AI skills tested by employers. I must admit, I think this number is actually low. I work with many different organizations. As I said, just this week, I've done workshops for four different organizations, and very few organizations actually test the skills of their employees. There's more and more companies doing training, hence why I'm so busy, which is awesome. But if you are looking at companies who are actually measuring the capabilities of their employees, this number is really, really low. So even the 64% of workers that never tested sounds really low to me. The number is probably in the 80% based on what my gut feeling tells me. Only 39% believe that their employer can effectively verify their AI capabilities, meaning it is not just that it's not being done, it's that the employees think that even if it is done, it'll be very hard to measure accurately. To that, I strongly disagree, but I think the employees themselves just don't have a clue how to measure their own capabilities, so that's why they think it's hard to do. Now, if you again drill down into the actual real numbers they have in the survey versus in the headline, 29% of Gen Zs explicitly lied about their AI skills. That's one in every three employees that are Gen Zs that specifically lied about their AI capabilities to their employers. that's a really high number, and that number, by the way, is significantly lower for Gen X and boomers But the real issue here is the level of anxiety, if you want, that employees from all ages, and again, younger, less experienced employees, stronger than others, 69% expect AI to automate parts of their role within the next 24 months. And the more interesting number is that 40% have personally observed AI tools performing their work. So that means that we have close to 50% of employees that were surveyed. Again, I'm gonna touch on how the survey was done in a minute and what I think about it, but let's stick to the findings for a minute. 40%, almost half of employees have observed AI perform some of their work on their own. They haven't read about it. They've actually seen it happen. That is a significantly higher number than I expected it to be. but it is showing you at least what people feel that is happening, right? Let's say the number is not real. Let's say the number is exaggerated. That's what people feel that is happening, and it is critical because that obviously drive, again, a lot of anxiety among employees, especially entry-level Gen Zs who just started in their workforce. Let's talk about for a minute who is this company and why they did the survey, and why I'm questioning at least some of it. GCheck is a background check company. Their sole purpose is to provide employers with better understanding of who they are going to hire, meaning they have a vested interest to show that people are going to lie to you because then you need their service even more. So I'm putting this out there so you guys know where the survey is coming from. In addition, the survey had about 1,500 people participate for a very short window of time and without compensating specifically for the percentages of the people in the population, meaning it's not a weighted outcome, and they had a lot more of, let's say, Gen Zs than other people participate, which obviously gives you distorted results. So I'm putting that out there so you know what's going on. but the bottom line is that people have serious anxiety with what AI is going to do to their jobs. Whether that's accurate or not doesn't really matter. Now let's look at two additional points to maybe bring this home. based on the recent statistics, data centers builds and operations have created 200,000 jobs in the US, and LinkedIn shows that 1.3 million new AI-related postings have appeared on LinkedIn. That a very big growth. So what is actually happening here, and are jobs going away, or are they going to grow? And I'll tell you what I think, and I've said that before. The first thing that I think is that, yes, new jobs are going to get created because of AI. We're already seeing it. That being said, the question that we need to ask ourselves is A) Whether it's going to create as many jobs as the jobs that it's going to take away. B), what is going to be the pace or the timeline in which these two things are going to happen? So if we compare the timeline in which jobs are lost, which is happening right now, versus when similar jobs are going to get created, which might be in the later future, there's a gap in between that might be a very serious bloodbath. The third thing that we need to think about, yes, data center created two hundred thousand jobs, But the people who lost their jobs do not have the skills that are required to do the jobs in these data centers. So here's where I think we're at. Do I believe AI will create new jobs? A hundred percent. Do I think it create as many jobs as the one it is going to take away? I personally don't think so. I think there are more jobs that are gonna be lost than jobs that are going to be created. I do agree, obviously, with the Chivons paradox, which basically says that the lower the cost is, the more the demand is. I'm seeing it myself. There are things that I couldn't sell before because the math didn't make sense, because I had to hire more people in order to do this. And now I don't have to hire more people in order to do this, and I can serve a larger audience, which means I can drive more value with less people, which means there's more demand for what I'm doing, which in the later run may drive me to add more people because the economics will make sense. However, again, in my very personal experience with myself and other companies, the math for now works in a way that you can grow with AI significantly more than most people think, which means you don't necessarily need to hire more people to service more clients, which means the paradox stops working once the actual job can be done by AI versus more people. So while the price is shrinking, the demand is growing. But then the question is: how do you serve that extra demand? Do you need more people to serve that extra demand or not? And my gut feeling tells me that, yes, you will need more people, but significantly less people than you would have needed in the past to do the same thing, which will still end with less growth than people anticipate, and that growth is required in order to compensate for the jobs that are being lost. The other aspect of this is training, retraining, and reskilling of people. Again, this is what I do for a living. I go from company to company, from organization to organization, and I help them train the employees and implement AI capabilities with the organization. This is something that allows me to see how hard it is to reskill people, and that is reskilling people in their existing role. You do things this way today, and we want you to do it that way tomorrow so you can be more efficient. this is not looking at somebody who has a specific profession and telling them to learn a completely new profession while using AI because these are the only jobs available. This is not easy to do, especially on a very large scale. And I think it will be required because yes, new jobs are gonna get created, but the people who are losing their jobs are not the people who can actually benefit from this. And the most important thing is the actual feeling of the individual and how they feel about the whole thing, and how does it impact their own wellbeing. Somebody who gets fired because of AI doesn't care about all of these statistics. He doesn't care, he or she, doesn't care about what may be coming as new jobs in two or three years. I love scuba diving, and I know I'm jumping to something completely unrelated, but you'll see how this connects in a minute. I love scuba diving, and I truly enjoy it, and I love scuba diving with sharks, and I know that sounds a little weird, but I think sharks are majestic, and they're incredible. And there's the statistics that a lot of people like to quote that's saying that more people get hit by lightning every year and die from it than people who get attacked by sharks. And yet people are a lot more terrified from sharks and getting into the water with sharks than people are worried about walking out when there's a storm outside. So this is more of a psychological thing. But what I say in addition to this, and this is how it will connect all together, if you're in the water with sharks and a shark attacks you, that statistics doesn't matter. The chance of getting hit by a shark goes from zero to 100 instantly, and then the stats don't matter. And it's the same thing here. If you are a person that just lost their job because of AI or AI washing, it doesn't really matter, you still lost your job. You're still unemployed. You still don't have money to spend on the things you want and/or maybe even the things you need, and you need to find a different job. And if you don't find that job quickly, your situation deteriorates, and if that happens on a very large scale, the economy as a whole deteriorates, which then impacts everything else. I think we're heading into a very problematic period from an economics perspective and from a job market perspective. I think there's going to be more job losses, and I think they're going to be in more than just the tech sector. And we're gonna end this episode today with talking about, again, a robot doing house cleaning, and you will see that wave is just around the corner of not just white-collar jobs, but also blue-collar jobs. So am I optimistic in the long run? Not sure. Am I pessimistic in the short run, so two to three years? From my perspective, I would say yes Now, the second topic I wanna dive into this week is how the AI race is changing and how what's important in AI is changing The first thing that I wanna quote is Stanford 2026 AI Index that documents that the top closed models now lead the top open weight models by just 3.3%. Now you can say yes, that changes. There's new models coming out. we just got, spoiler alert, we just got Claude Opus 3.8. We are expecting probably the same thing from OpenAI in the near future. So whatever they're gonna call it, maybe GPT 5.6 or whatever it is going to be. But still the gap from the open source models is shrinking If you look at the top 10 models on the regular chat text arena, the spread is with Claude Opus 4.6 Thinking on the top with 1503, and on number 10 you have GPT 5.5, the regular version, not the high version, with 1474. That's a 50-point spread over a 1500 score. That is negligible. And in there you have models like GLM 5.1 and Grok 4.2 and Qwen 3.7 and Grok 4.2 Beta, and even Muse Spark by Meta. All of them are very, very close together from a capabilities spread perspective. And yes, they vary on coding and they vary on image generation, but if even on those areas, the spreads are not that big. And by the way, even where the spreads are bigger, and there are cases like in the agentic side where the closed models are significantly better than the open source models, the open source models are good enough. They can do most work that is required in the knowledge work today really well or good enough in order to replace people or at least dramatically increase their capabilities and their throughput. Now, if you wanna look at it not from a quantitative perspective of the different benchmarks, but from a qualitative perspective, Adina Jacub from Hugging Face related to the latest release from Minimax and said their Really solid work on mixture of agents, efficiency, and agent-oriented design, and she praised on how they are now coming very, very close and on some cases doing better than OpenAI and Anthropic on a few different things after the gap between them was significantly higher not too long ago. So she continued and said, "Beyond the benchmarks, they've done some really solid work," which tells you that they are delivering a very good model, and she was specifically talking about minimax M3, which was just released, which is doing some really cool things which we're gonna talk about in a minute Now the second aspect of this conversation, beyond the fact that the models are getting closer and closer together, and like I said, most of them, including the open source ones, have the capability to do most jobs in a pretty solid way, even if some doing it better than others. There is the budget question of this, and the costs companies are paying right now are going through the roof. The most viral announcement this week that caught fire and appeared on almost every publication you can imagine was that a company that wasn't named got to $500 million of spend in just one month to Anthropic because they did not put any caps or limits on what employees can use in that. But to be fair, that wasn't confirmed by anyone. This appeared as one line in a Axiom newsletter that then was quoted by everybody else. But the name of the company wasn't published, an invoice wasn't shared. Anthropic didn't say anything about this. And tried to do the math in order to see if that even makes sense So let's do the math together so we'll understand how insane that number is. Let's say the average price they paid for the output tokens is $10. There's a big variety. It can be $5, it can be 60 cents, it can be $25, but let's take $10 as an average just to make the math easy. If we take $10 as the average to get to trillion tokens in a single month. If you want to understand what that means, that is 19 million tokens every single second, nonstop, 24/7 for the entire month Now, if you think about when you just run Claude regularly, it generates about 50 to 80 tokens every single second. It means you gotta run hundreds of thousands of sessions every single minute across multiple users nonstop for an entire month to get to that number. Now, if you want a different benchmark, Uber, which we're gonna talk about in a minute, said that their heavy users are using about $1,000 to $2,000 per month in tokens. To hit 500 million even at the $2,000 mark, you need 250,000 engineers running applications Again, these numbers just don't add up. They don't make any sense. I'll be really surprised if that number is real. And I think if that number was real, we would have known by now who the company is or exactly what happened and stuff like that. None of that became available. Again, it was one line in Axiom that everybody jumped on and shared. That being said, going back to Uber Uber admitted that they burned through their entire annual budget to Claude by the end of April and they shared that Claude code adoption jumped from 32% of employees to 84% of employees through its 5,000 engineers in just a few months. The per engineer monthly cost now varies between $500 on the lower end to $2,000 on the higher end per month. Their CTO said, "The budget I thought I would need is blown away already." But in addition, what they're finding is that it's very, very hard to quantify the value of the tokens that they're using. So Andrew McDonald, their COO said, "It is very hard to draw a line between one of those stats and okay, now we're actually producing 25% more useful consumer features." We also talked about the fact just a couple of weeks ago that Microsoft is canceling internal Claude Code licenses, pushing everybody to use their own GitHub Copilot CLI, and that the deadline for all of that to cancel all of the Claude capabilities is June 30th. And that, again, is obvious because it looks bad when you are using the competitor's tool instead of your own internal tool. But in addition, we talked about the fact that they're about to close their fiscal year, and hence, cutting cost could play a very big role in showing efficiencies. And so again, they share that their per engineer cost is five hundred dollars to two thousand for Claude Code's tokens. So what does that tell us? It tells us that companies are spending more and more money on tokens without necessarily knowing exactly what they're spending it on what the results that it's generating, which is very, very hard to measure. And that is, by the way, happening while the actual token cost, the cost per token, is dropping dramatically. So dropped over 80% in the past year. So we're talking about almost an order of magnitude decrease in the cost per token because the technology is getting better, but with cost per usage are dramatically increasing, mostly because more and more employees and more and more organizations are jumping into the agentic universe, and the agents are consuming crazy amount of tokens. So despite the fact that tokens are becoming cheaper, the overall cost is growing very high, very, very quickly. So this is the current situation, and this is why I think the conversation is changing, and it's changing fast from being a conversation about which model is better to a conversation about what value we can get and how fast we can get it and how easy we can get it, and that becomes more important than this benchmark or that benchmark that this model or that model can get a higher score on. So we got two interesting examples about this week. Google just released Gemma 4. Gemma 4 has a capability that's called multi-token prediction or MTP, which drives three times faster results than the previous Gemma model at the same quality. So how does MTP works? MTP has two separate components. One of it is called drafters, which basically try to speculate across multiple tokens at the same time, trying to drive the speed higher. But the second layer is the primary Gemma 4 model that verifies the suggestions and in parallel to the generation of them, what is overall, as I mentioned, generating three times faster token generation. And it can do this across all the different model sizes that Gemma 4 comes in. Now, if you look at what's happening recently in the open space model, you would see that in the past two months, roughly every open source model has included such capability in it, so MTP is becoming a standard which is allowing AI to provide results significantly faster without changes in the hardware and without changes in the actual results from a quality perspective. Another approach to this is Minimax-M3 that we talked about earlier. Minimax M-M3 is achieving 15.6% faster decoding for a one million token context window. They are using a technology that they called Minimax Sparse Attention, or MSA. But it does something similar where it performs block level selection on real uncompressed keys, meaning they are also running things in parallel versus token by token, which is how AI worked until very recently, and they're achieving also 99.7 faster pre-filling at the full one million context window. What does all of this technical mumbo jumbo means? It means that M3 achieves better results than M2 while reducing the compute cost by 80% and running faster To put things in perspective, Minimax M3, which is not at the frontier level, but it's just slightly behind, cost $0.3, so 30 cents for every million entry tokens and $1.2 for million output tokens And that model will be open source by the time you listen to this episode To put things in perspective, compared to the 30 cents of input tokens, Claude 4.8 costs $5 for million input tokens. And if you compare the output tokens You'll be paying $25 on Claude versus $1.2 on MB by Mistral. That is almost 25X the cost for some benefit in the value, Definitely not 25x value. Now, on the same topic, we heard Perplexity's CEO, Arvind Srinivas, quoting his new view on who's going to win in the AI race based on the value per token, or specifically the way that he phrased it is token value per watt per user. And he said that this parameter, token value per watt per user, is going to be, and I'm quoting, "Whoever is able to maximize this particular objective really will by balancing accuracy, latency, cost, privacy, and intelligence all together, they're going to win. That's what's going to win in the long run." And I agree with him. It is very, very clear that currently all the leading models are just good enough. What is going to make the difference Are three things, how easy it is to get to the outcome, and that combines all the things he said, like the ease of use, the latency, the cost, the privacy, the stuff like that, together with how cheap and quickly can I get it. This is going to be more important by a very big spread than the final capability of the model itself, because the not leading models are also going to be good enough, and I believe we are already there. Which leads us to the last component of what is going to make a difference, and you heard me say that before, I'm just putting everything together, which is the harness. So what the hell is a harness? The harness is how you are actually going to work with the AI model. If you think about it, the AI model is the engine and the harness is the car, right? And you're going to pick a car not necessarily just because of its engine, just like you do today, right? You don't really care which engine's in the car. You want the specific car because it serves you better, you can afford it better, it does what you want it in a better way, and it doesn't have to have the best engine inside. It just have to serve your needs, and it is exactly the same thing with the harness and the model. The model is the engine. It's the technology that runs the results and the intelligence inside of it. But how you can approach it, how you use it makes a very big difference. Better than nothing, but has lesser capabilities from a model perspective? Absolutely not. On the benchmark, GPT 5.5 is winning on most cases. However, the ease of use to make the best out of this technology in the Claude CoWork environment is just a lot easier for me to get the results I'm looking for than doing the same exact thing in the ChatGPT environment. Are they making changes in order to catch up? 100%. If you're now on the enterprise level, there's an agent capability built into the app. They are going to release Codex into the ChatGPT app, just like you have Claude Code inside the Claude app. So they're closing the gap. But as of right now, and in the last six months, they were ahead, not by the model's capabilities, even though in some cases it is, but mostly by the ability to make value out of these capabilities, which is what really matters. Combine that with the fact that the leading software companies are doing everything they can to embed AI into their platforms. That is true for Salesforce. The most out of the platform you're working with is what's going to make a difference, not necessarily caring about which model is running behind the scenes. so far we understand that the speed makes a very big difference, the cost make a very big difference, the harness makes a very big difference. And then we have the last component, which is the orchestration across different platforms or different compute capabilities is also going to make a difference because it is going to combine a lot of the things that we talked about before. So we see more and more companies that are enabling AI to run across different platforms. If you have tried Perplexity's computer capabilities, which are really, really awesome and really amazing, they use some of the capabilities of your operating system, running some of it locally, and then automatically decides what to actually send to the cloud or like Arvind Srinivas, the CEO of Perplexity, said, "The data center is coming to your laptop." Another good example that we're expecting in the next couple of weeks is Apple is supposed to finally introduce the new Siri that is powered by Google Gemini, but it is also going to use on-device capabilities and jump to the cloud just when it needs to do so. Microsoft just launched a new MAI family of products which we're gonna talk about in a minute in the next topic, in their latest event, and they are able to run locally to some extent on your computer hardware without going to the cloud. And then we're going to see the same thing for new devices. We heard a lot of new rumors on devices as well. so we're going to see more and more models that know how to run on your wearable device, whether it's a watch or glasses or a pendant or whatever it is that you're gonna be wearing, and can run on your local computer, PC, et cetera, and can run on the cloud. And jumping back and forth between them will give more speed, more privacy, and cost savings that can be transferred to I can now create or gain more value without paying the high price or the latency that we are used to or willing to accept right now and will not be willing to accept in, I think, the very near future So what does that mean if you are running a business and you need to decide which models to use? One thing that we didn't talk about is what the companies are trying to do in order to lock you in to their environment. As an example, if you want to use any of the enterprise level platforms, they demand you commit for a year, and they demand for at least fifty licenses to be included in that full year. So you're already locked in for a full year for at least some of your employees. That's one lock-in. The other lock-in is obviously the integrations to different tools and different systems that you're going to build. And over there it is your choice. You need to be smart about this and to try to build a layer that integrates your data into the AI as generic as possible. Because if you build it specifically for ChatGPT or specifically for Claude or specifically for something else, you will be locked in forever instead of just for the full year you're committing to, versus being able to switch whenever you want with relatively little effort in order to gain the benefits from another platform that is now cheaper, faster, more efficient, better harnessed and so on. So why am I sharing all of this with you? This makes a very, very big difference how to plan your AI tools, how to plan your budget, and how to plan your training. Because if training your employees on a specific platform versus generic use of AI, and then you can use whatever tools, then again, you are increasing your lock-in to one platform over the other. All tools out there today are good enough for most white-collar knowledge work if you train your people and you have the right connectors and the right data access to clean data in a safe way. What I just said in one sentence is very not obvious. Again, neither the training nor the access to clean data in a safe way. But if you do this path connected to we are committed just to OpenAI or just Anthropic or just Gemini or just Grok, doesn't matter, you will see that later on you will pay a higher price than you need to. And if you think about it right now and you try to build it as generic as possible, you'll be able to gain significantly more value per dollar in the future. And like I said, I think the race is changing. I think the race is going from we need better models to we need easier, faster, cheaper access, safer to generating real value from these models. and I think the race, while it is definitely on, if you just look at the releases from Anthropic the past six months, they released, I don't know, like six different new models. And I don't see that stopping, but I think the bigger difference is going to be how it integrates to our real life, our real business, and do this quicker, cheaper, faster, and safer for the users. And as users, we need to learn how to switch and make it as painless as possible so we can enjoy these benefits. Now let's jump to rapid fire of this week. The first rapid fire item is Microsoft held their Build conference this week. They have announced seven new in-house MAI models. MAI stands for Microsoft AI. So these models span across more or less anything you can imagine. It includes MAI Image 2.5 plus a flash variant of the same thing, speech to text, which is called Transcribe 1.5, reasoning, which is called Thinking 1, text to speech, which is called Voice 2, and coding, MAI Code 1 Flash, which is already live in Copilot and VS Code. So this is a really wide range of tools, and you can imagine that this becomes a very important part of their strategy as they're trying to break away from their dependency on OpenAI. So we've seen the few first steps as they partner with Claude for more and more things, but now building their own models is obviously gonna give them more freedom. They also has upgraded Microsoft IQ, which is a context layer for agents, and it's now generally available. So that was available before, but just to select companies, and now it is available across GitHub Copilot Foundry and Copilot Studio, and it has three different layers. It has Work IQ, which basically knows everything in your organization and how it actually works, email, meetings, docs, whatever you want to connect to it. Fabric IQ, which is a structured business data and work, and Web IQ, which is fast AI first web grounding to verify facts and so on. And this is going to be publicly available in the next couple of weeks. So Work IQ APIs go to general availability on June 16th. They also introduced Microsoft Scalp, which is a proactive personal work agent that is built on top of Work IQ and handles everything from meeting prep to scheduling colleagues to, routine tasks without being even in a Microsoft environment and can help you do your job or anticipate what you need and help you with that. They also introduced Project Solara, which are agent-first devices. And interestingly, they're built on Android and not Windows, so they're talking about agents that will run locally on the device. We heard that from the rumors about where OpenAI is going with their device. But the idea is that instead of running apps on devices, we will run agents on devices to do the things that we need them to do, and Microsoft are going to be one of the players in that game, again, running specifically on Android. They also clearly pushing Foundry as an agent factory. So they introduced Agent Framework 1.0, which is now generally available. And the idea is you can now host and run long-running agents and procedural capabilities inside of Foundry native connected to everything else Microsoft, including publishing it to Teams and 365 Copilot. So it will run in Foundry, but you will have access to it in your regular day-to-day tools. They also introduced local AI on Windows. So this is a small, fast model that can run entirely on an NPU for email summaries formatting, basically basic tasks that doesn't need the cloud at all, going back to the topic we just talked about, uh, before in where this is all going, and this is going to be available on the new versions of devices and models that they're sharing. And they're upgrading their governance and money layer. They're calling it Agent 365 that allows companies to govern agents and Copilot credits to meter their work. So going back again to the topics we started this conversation with on how do you manage your organization? How do you manage how many credits are people using? Where are they using it? What are the value that it's generating? It's starting to get built into the platforms themselves. The next rapid fire that I already hinted to, Anthropic gave us Opus 4.8 it is their most capable model that they've ever released. Not surprising. It has a one million tokens context window, and it is designed, again, not surprising, for advanced coding, AI agents, and enterprise workflows And also it is scoring higher on most benchmarks than Opus 4.7 and GPT 5.5. Again, not surprising, otherwise I don't think Anthropic would have released it. The model automatically adjusts computational effort between different tasks depending on the complexity of different steps of the different tasks. So it is going to think harder on more complex problems, and it's going to think less on less sophisticated or needing problems, which is going to save you money as a user and give you faster results when it doesn't need to think very hard. It is priced exactly the same as Opus 4.7, so five dollars input tokens and twenty-five dollars for a million output tokens. Also similar to before, caching prompts, which helps a lot, is going to save you up to 90%. If you do that, you need to remember these expire, so you can only use them during a specific session. Once it's over, if 30 minutes passed, you lose those caching unless you use really sophisticated tools out there that enables you to save those in different ways. And as I hinted before, this is a crazy release cycle by Anthropic. If you think about it, they released Opus 4.1 in August of 2025, and then in November we got 4.5, in February we got 4.6, in April we got 4.7, and now we got 4.8, and they seem to be releasing them faster and gaining more and more capabilities. But as I mentioned previously, that doesn't necessarily matter. But that being said, if you look at benchmarks that actually may be connected to real life, such as GDPVal, which is trying to evaluate actual work that needs to be done across different aspects of the organization, Opus 4.8 scored 1890 ELO points, 137 points above Opus 4.7, 121 points ahead of GPT 5.5 Extra High. So very capable model, another very capable model, I should say, from Claude. But the interesting thing is, connecting it back to our previous topic, it is doing this while consuming 35% fewer output tokens, which means you're gonna pay less, and you're gonna get the results faster than you did in Opus 4.7, which is great for all of us. On the other side of this race, we have OpenAI, who just announced on June 4th a new version of its memory. It's called Dreaming 3.0, which is supposed to provide much better, more effective memory inside of the OpenAI models. It is supposed to be a completely new infrastructure to how their memory works and allows to provide better, cheaper. They're also going to roll it out later on to all free users. Now, as part of this release, they shared something I found interesting, which is how they evaluate memory success, which is on three different dimen-dimension: following user preferences and constraints consistent by staying current as time passes, like updating. I've mentioned two weeks before the party and now it's three days before the party, and you'll be aware of that and we'll be able to provide the right guidance. OpenAI also announced that they are going to be integrating Codex into the OpenAI app. This haven't happened yet, but it is going to happen in the next few weeks. On my application, it's still not showing up, but this rollout is accompanied by new capabilities, and Codex now has Sites, which I find really exciting. So what is Sites? It allows you to deploy whatever you develop from Codex directly inside the OpenAI environment. So if you've been using any of the vibe coding tools, you know that there's two separate approaches. The more, if you want professional tools like Codex or VS Code, require you to get a deployment environment. You need somewhere to deploy the code that you generate, the same way that professional software companies are doing it. I use mostly Railway and Google Cloud, but I also have some stuff on AWS. So you need to learn how to create this deployment mechanism. There are tools on the other side, like Replit and like base44, that allow you to deploy on their environments, meaning you need to know nothing about how to deploy software and so on. You just click Deploy, and then it's available to whoever you want to make it available to. And now this functionality is going to be available inside of your ChatGPT application Because Codex is gonna be available on your ChatGPT application, and Codex now has sites, which is the ability to do the deployment straight from the ChatGPT environment. This is a shot directly towards competitors like Replit and like Base44 and definitely for Anthropic. Anthropic has the capability to share on a limited way. I think this new functionality from OpenAI is gonna push Anthropic to do the same thing. It will be very interesting to see how much that is going to cost over time from a hosting and usage perspective, but the functionality is definitely helpful from small applications to dashboards and stuff like that you can now generate in minutes and then deploy and make it accessible to other people in your organization and beyond. As part of this release, they also released six new agent plugins that are covering different aspects of work. They include data analytics, creative production, sales, product design, public equity investing, and banking. So each and every one of these gettable inside OpenAI environment. We're seeing very similar things from Anthropic. We're gonna talk about this in a minute What they also shared, which is not surprising because the same exact thing happened on the Anthropic side, is that currently, after their crazy spike in usage that they've seen since the launch of Codex, 20%-ish of Codex users are not software engineers, and they're using it to do other things, which hence makes a lot of sense that they're adding this to the regular OpenAI application and that they're adding sites to allow people to deploy what they're actually building. They also announced that their latest models, so GPT-5.5 and 5.4 and Codex software capability for engineering, for engineers are now generally available on Amazon AWS via Amazon Bedrock. This is a interesting strategic move that really is a great partnership that helps both parties because from OpenAI's perspective, they're breaking away from the exclusive agreement they had with Microsoft, which was shared with you a few weeks ago. So there's no exclusive agreement anymore. But in parallel, this is a great opportunity for Amazon because as part of this agreement, OpenAI are going to use their Trainium chips instead of NVIDIA. So Amazon is getting a huge growth to its Trainium chip deployment, and OpenAI is getting another delivery channel through AWS Bedrock. Both companies are winning, and us as users are winning as well because if you're on AWS, you can now have access to these models that were not accessible before. We talked about launching new capabilities, and we said that OpenAI released their capabilities while Anthropic also released a lot of new plugins and new tools. So they just announced they're dramatically expanding their Claude for Legal capabilities. So beyond the original 12 plugins that they released, there are now 90 specialized AI agents designed for specific granular legal work. So Claude for Legal now provides many different aspects of law work Trying to cover end-to-end workflow of a legal person, including vendor agreement reviewer, responder, claim chart builder, and so on and so forth. We're not gonna dive into this. Now, in addition, Claude for Legal includes over 20 new MCP connectors to more or less any tool you wanna use in the legal work, such as DocuSign and Ironclad and iManage and NetDocuments and LexisNexis and Thomson Reuters and many, many others, which means you can do more and more professional legal work, ei-either as a law firm or as internal legal teams in companies with Claude versus going to the professional tools like Harvey, which may be more capable. But again, is that more capable worth the extra cost? Is question that still needs to be answered. It is very, very clear that both these companies are going hardcore after specific tasks in specific organizations, and they're gonna keep on pushing in that direction because, again, that's next frontier. We talked about the speed and cost to you out of the box, the capabilities to do the work you need to do to get the value you need paying them versus somebody else, and overall saving you money because you don't need two or three or four services now a few interesting news about hardware side of AI. Meta is reportedly developing an AI-powered pendant, and it is testing it for a next year release. This is following their acquisition of the company called Limitless, which is just acquired recently, and this is based on an internal memo from that was published by the Information. Now This initiative is a part of a broader strategy by Meta to expand its AI glasses lineup and introduce what they call wearable for work, which is going to include a business subscription. In order to be able to use this, you will need to pay them an annual subscription in order to enjoy the benefits of this. And this is trying to compensate for the crazy losses they currently have at the Reality Labs division. So Reality Labs has lost four point, $4.03 billion just in the first quarter of 2026, accumulating to over $80 billion in losses since its inception and 19 billion just in 2025. So the idea here, like we've seen OpenAI and Anthropic, push from the consumer market to the business professional market and then be able to charge higher prices and per seat if you want licenses, and that's the direction that they are trying to go. So apparently it's going to be a pendant that will be used mostly for business use cases, and in order to use it, you will need to pay a monthly or an annual fee to enjoy the benefits. I anticipated this before, and I'm gonna say this again, I think sometime in the not-too-far future, we are all going to be wearing AI devices because it makes perfect sense, and that raises a lot of questions from privacy, relationships, security, data access, and a lot of other stuff that I don't think anybody has answers for. But I think the hardware will start being more and more available and more and more worn and used by more and more people, and we'll have to figure that out as we go. And another interesting, cool, scary, weird, call it whatever you wanna call it piece of news from this week is that San Francisco-based startup called Gatsby has completed the first consumer home cleaning by a humanoid robot in the United States. So if you're wondering what the hell this is, Gatsby provides a robot that you can book from an iOS app for $150 flat fee per apartment in San Francisco. So it doesn't matter how big your apartment, a robot is going to show up at your doorstep and invest X amount of time in cleaning and doing everything you need as far as cleaning. That includes dishes, surfaces, floors, bed making, laundry, folding, all without a human person there watching it. It completed the first task in about three hours. What does that tell you? not much. I'm not sure how well the cleaning went. There are also a lot of open questions, including privacy or what happens if it breaks something or harms someone while it is doing what it is doing. It does have remote access of humans, so if it gets stuck in anything complex, there's a human remotely who can access and see what the robot is doing and control the robot, which raises even more concerns from a privacy and control perspective. But the bottom line is very, very simple. This is a robot that just did a house cleaning for a fee that is at the very, very low end of what humans charge in the San Francisco area, which is about one hundred and fifty to three hundred dollars to do the same exact work. So while this robot may not be perfect, it is the canary in the coal mine that tells you what is happening. We are going to see more and more robots take blue-collar work and do it at a rate that humans will not be able to sustain or not be able to sustain themselves if that's what they get paid, That's gonna push a lot of uncertainty into the lives of people. This is the kind of work it is going to be doing. if this time it is just cleaning house, the next thing is gonna be baristas and different service providers and cashiers or tellers or customer service at railway stations and so on and so forth. We are going to see more and more of these showing up in the next few years, which is going to drive more people out of their existing jobs, which as I mentioned, I'm not very optimistic on the white collar knowledge work. This is just the next wave, and it is coming, and even this is not gonna grow very, very fast. It will grow in the next three to five years to scales that will start make a difference Before I let you go, I want to remind you of our courses. The next multi-agent orchestration course is starting at the beginning of August. We sold out of every course between May and the end of July, so this is definitely something that you want to do if you want to understand where the world is going, how to build multiple agents, and how to make them work together to automate more or less anything you can imagine in your business. I'm also doing this as workshops for companies, so if you're in a leadership position in a company and you want to teach this to the relevant group of employees in your company, I will gladly talk to you about this. You can grab time on my calendar with the link on the show notes, or just find me on LinkedIn and send me a message. Both ways are perfectly fine. We will be back on Tuesday with another how-to episode that will teach you how to use AI in your business right now. And until then, have an amazing rest of your weekend.