1
00:00:00,000 --> 00:00:01,720
Everyone keeps asking the wrong question,

2
00:00:01,720 --> 00:00:03,920
is MAI one better than 5/4?

3
00:00:03,920 --> 00:00:05,640
Which one should we standardize on?

4
00:00:05,640 --> 00:00:07,140
That question was already outdated,

5
00:00:07,140 --> 00:00:09,360
the moment Microsoft shipped both of them on purpose

6
00:00:09,360 --> 00:00:10,160
in the same month.

7
00:00:10,160 --> 00:00:11,800
Here is the tension nobody is naming.

8
00:00:11,800 --> 00:00:13,960
Microsoft just released two completely different types

9
00:00:13,960 --> 00:00:15,920
of intelligence, not a bigger version

10
00:00:15,920 --> 00:00:17,640
and a smaller version of the same thing,

11
00:00:17,640 --> 00:00:19,720
two different organs built for two different jobs.

12
00:00:19,720 --> 00:00:21,120
Most of the coverage treats them

13
00:00:21,120 --> 00:00:23,680
like they are fighting for the same spot on the podium.

14
00:00:23,680 --> 00:00:26,240
But they are not MI1 and 5/4 are not rivals,

15
00:00:26,240 --> 00:00:29,480
they are not competing for the same job inside your stack.

16
00:00:29,480 --> 00:00:31,560
By the end of this, you are going to understand

17
00:00:31,560 --> 00:00:34,040
the actual architecture Microsoft is building toward,

18
00:00:34,040 --> 00:00:37,280
not a feature list, not a benchmark chart, the real system.

19
00:00:37,280 --> 00:00:39,400
If you are the kind of person who has to design this stuff

20
00:00:39,400 --> 00:00:41,160
for a living, subscribe now.

21
00:00:41,160 --> 00:00:42,600
Because this split is going to define

22
00:00:42,600 --> 00:00:45,360
how M365 admins and architects build systems

23
00:00:45,360 --> 00:00:46,960
for the next five years.

24
00:00:46,960 --> 00:00:48,800
The question everyone is getting wrong.

25
00:00:48,800 --> 00:00:51,120
So let's start with how most people are covering this.

26
00:00:51,120 --> 00:00:52,840
It explains why so many teams are about

27
00:00:52,840 --> 00:00:54,440
to make expensive mistakes.

28
00:00:54,440 --> 00:00:56,320
Watch any breakdown of MI1 and 5/4

29
00:00:56,320 --> 00:00:58,160
and it turns into a leaderboard battle,

30
00:00:58,160 --> 00:01:00,040
which one scores higher on this benchmark,

31
00:01:00,040 --> 00:01:02,920
which one is faster, which one is smarter per parameter,

32
00:01:02,920 --> 00:01:03,920
bigger versus smaller.

33
00:01:03,920 --> 00:01:06,360
Like it is a heavyweight fight and somebody has to lose.

34
00:01:06,360 --> 00:01:08,480
That framing comes from consumer AI thinking.

35
00:01:08,480 --> 00:01:10,560
It is the same instinct that made everyone argue

36
00:01:10,560 --> 00:01:12,800
about which chatbot writes better emails.

37
00:01:12,800 --> 00:01:14,520
And look, that instinct isn't stupid,

38
00:01:14,520 --> 00:01:16,520
it is just built for the wrong problem.

39
00:01:16,520 --> 00:01:19,360
When you are a consumer picking one assistant for one phone,

40
00:01:19,360 --> 00:01:21,960
which model is better is actually the right question.

41
00:01:21,960 --> 00:01:23,160
You only get to pick one,

42
00:01:23,160 --> 00:01:25,280
but enterprise system design does not work that way.

43
00:01:25,280 --> 00:01:26,120
It never did.

44
00:01:26,120 --> 00:01:27,840
Nobody runs one server for every workload.

45
00:01:27,840 --> 00:01:30,160
Nobody uses one storage tier for every file.

46
00:01:30,160 --> 00:01:32,080
So why would anybody assume there is supposed

47
00:01:32,080 --> 00:01:33,880
to be one model for every request?

48
00:01:33,880 --> 00:01:35,120
This is worth naming directly.

49
00:01:35,120 --> 00:01:37,320
It is the assumption sitting underneath almost everything

50
00:01:37,320 --> 00:01:38,600
written about this launch.

51
00:01:38,600 --> 00:01:40,120
Call it the model thinking.

52
00:01:40,120 --> 00:01:42,560
The old belief that one system has to win.

53
00:01:42,560 --> 00:01:44,520
The idea that intelligence is a single throne

54
00:01:44,520 --> 00:01:45,920
and somebody is sitting on it.

55
00:01:45,920 --> 00:01:48,840
That belief made sense when there was basically one option,

56
00:01:48,840 --> 00:01:50,440
rent GPT or rent something else,

57
00:01:50,440 --> 00:01:52,280
but that is not the world Microsoft just built.

58
00:01:52,280 --> 00:01:53,280
Here is the reframe.

59
00:01:53,280 --> 00:01:56,560
Microsoft is not choosing a winner between MI1 and 5/4.

60
00:01:56,560 --> 00:01:58,040
It is building a division of labor.

61
00:01:58,040 --> 00:01:59,080
Once you see it that way,

62
00:01:59,080 --> 00:02:00,880
the leaderboard question falls apart

63
00:02:00,880 --> 00:02:02,440
because you are not supposed to pick one.

64
00:02:02,440 --> 00:02:04,240
You are supposed to know when to use which.

65
00:02:04,240 --> 00:02:05,480
So let's set up the two words

66
00:02:05,480 --> 00:02:07,840
that are going to anchor everything from here forward.

67
00:02:07,840 --> 00:02:09,360
Reason and runtime.

68
00:02:09,360 --> 00:02:11,640
Reason is the deep layer, the part that plans,

69
00:02:11,640 --> 00:02:12,840
the part that weighs options,

70
00:02:12,840 --> 00:02:14,600
it figures out what should actually happen next

71
00:02:14,600 --> 00:02:16,640
in a complicated multi-step problem.

72
00:02:16,640 --> 00:02:18,880
Runtime is the fast layer, the part that executes,

73
00:02:18,880 --> 00:02:20,640
the part that responds instantly.

74
00:02:20,640 --> 00:02:23,000
It lives close to wherever the work is actually happening.

75
00:02:23,000 --> 00:02:24,520
MI1 sits in the reason slot,

76
00:02:24,520 --> 00:02:26,200
5/4 sits in the runtime slot.

77
00:02:26,200 --> 00:02:27,480
Once you frame it that way,

78
00:02:27,480 --> 00:02:30,080
comparing their benchmark scores directly is a mistake.

79
00:02:30,080 --> 00:02:32,440
It is like comparing a company's strategy team

80
00:02:32,440 --> 00:02:33,840
to its shipping department

81
00:02:33,840 --> 00:02:35,560
and asking which one is better.

82
00:02:35,560 --> 00:02:36,600
Wrong question.

83
00:02:36,600 --> 00:02:38,080
They are not measured against each other.

84
00:02:38,080 --> 00:02:40,320
They are measured against the job they were built for.

85
00:02:40,320 --> 00:02:42,640
That is why benchmarks alone cannot tell you what to do

86
00:02:42,640 --> 00:02:44,280
with either of these models.

87
00:02:44,280 --> 00:02:45,440
A number on a leaderboard

88
00:02:45,440 --> 00:02:48,320
does not tell you where a model belongs in your architecture.

89
00:02:48,320 --> 00:02:49,640
What tells you that is understanding

90
00:02:49,640 --> 00:02:52,080
what actually happened at Bill 2026?

91
00:02:52,080 --> 00:02:54,600
That is when Microsoft laid this whole strategy out in public

92
00:02:54,600 --> 00:02:57,440
and what actually happened at Bill 2026.

93
00:02:57,440 --> 00:02:59,200
So here is what actually got announced.

94
00:02:59,200 --> 00:03:01,120
The framing matters more than the coverage.

95
00:03:01,120 --> 00:03:02,960
Mustafa Suleiman walked on stage

96
00:03:02,960 --> 00:03:05,840
and unveiled seven new MI models in one sitting,

97
00:03:05,840 --> 00:03:07,960
not one flagship, not a single headline model.

98
00:03:07,960 --> 00:03:09,400
Seven spanning reasoning, coding,

99
00:03:09,400 --> 00:03:11,240
image generation, voice and transcription

100
00:03:11,240 --> 00:03:13,800
all released together, all trained by Microsoft's own team

101
00:03:13,800 --> 00:03:14,760
from scratch.

102
00:03:14,760 --> 00:03:16,960
And alongside that, 5/4 kept expanding too.

103
00:03:16,960 --> 00:03:18,920
More variance, more sizes,

104
00:03:18,920 --> 00:03:20,520
still shipping on its own track,

105
00:03:20,520 --> 00:03:23,160
two families growing at the same time in the same event.

106
00:03:23,160 --> 00:03:25,280
Microsoft didn't call this a product launch.

107
00:03:25,280 --> 00:03:27,800
They called it humanist superintelligence.

108
00:03:27,800 --> 00:03:30,440
State of the art capability across modalities,

109
00:03:30,440 --> 00:03:33,280
but explicitly designed to serve people and organizations

110
00:03:33,280 --> 00:03:34,120
not replace them.

111
00:03:34,120 --> 00:03:36,600
You can take that branding with a grain of salt if you want.

112
00:03:36,600 --> 00:03:37,560
That's fair.

113
00:03:37,560 --> 00:03:39,840
But underneath the phrase is a real claim.

114
00:03:39,840 --> 00:03:43,120
These models are meant to work across text, code, image, voice,

115
00:03:43,120 --> 00:03:45,320
and reasoning as one coordinated system,

116
00:03:45,320 --> 00:03:46,960
not a scattered side projects.

117
00:03:46,960 --> 00:03:48,920
Now here is the part that actually explains why

118
00:03:48,920 --> 00:03:51,240
this happened right now in this exact window

119
00:03:51,240 --> 00:03:52,320
and not two years ago.

120
00:03:52,320 --> 00:03:54,840
Until late 2025, Microsoft literally

121
00:03:54,840 --> 00:03:56,320
could not build frontier models.

122
00:03:56,320 --> 00:03:57,720
That wasn't a strategy choice.

123
00:03:57,720 --> 00:03:59,320
It was contractual.

124
00:03:59,320 --> 00:04:02,520
The original OpenAI partnership signed back in 2019,

125
00:04:02,520 --> 00:04:05,200
restricted Microsoft from independently pursuing

126
00:04:05,200 --> 00:04:08,320
frontier level AI or superintelligence on its own.

127
00:04:08,320 --> 00:04:10,880
So for years, Microsoft's entire AI story

128
00:04:10,880 --> 00:04:14,840
was really OpenAI's AI story, wearing a Microsoft badge.

129
00:04:14,840 --> 00:04:17,360
Every leap in capability ran through somebody else's lab

130
00:04:17,360 --> 00:04:17,760
first.

131
00:04:17,760 --> 00:04:20,520
That changed when the contract got renegotiated.

132
00:04:20,520 --> 00:04:22,040
And once that restriction lifted,

133
00:04:22,040 --> 00:04:24,360
Microsoft didn't ease into building its own models.

134
00:04:24,360 --> 00:04:25,200
It sprinted.

135
00:04:25,200 --> 00:04:27,280
Think about what that actually means strategically.

136
00:04:27,280 --> 00:04:29,440
In under a year, Microsoft went from being

137
00:04:29,440 --> 00:04:31,760
a distributor of intelligence to being a builder of it.

138
00:04:31,760 --> 00:04:33,560
Those are completely different businesses.

139
00:04:33,560 --> 00:04:36,120
A distributor's job is picking the best thing

140
00:04:36,120 --> 00:04:38,480
off somebody else's shelf and packaging it well.

141
00:04:38,480 --> 00:04:41,120
A builder's job is deciding what gets made in the first place,

142
00:04:41,120 --> 00:04:43,720
what trade-offs get baked in, what the model is actually for.

143
00:04:43,720 --> 00:04:46,640
Microsoft crossed that line in months, not years.

144
00:04:46,640 --> 00:04:48,760
And normally, when you hear seven new models

145
00:04:48,760 --> 00:04:50,240
shipped in one announcement, you

146
00:04:50,240 --> 00:04:53,200
assume some giant sprawling org chart behind it.

147
00:04:53,200 --> 00:04:55,720
Thousands of researchers, endless review cycles.

148
00:04:55,720 --> 00:04:57,240
The usual big company machinery.

149
00:04:57,240 --> 00:04:58,360
That's not what happened here.

150
00:04:58,360 --> 00:05:00,440
Suleiman described teams of around 10 people

151
00:05:00,440 --> 00:05:01,760
building these models.

152
00:05:01,760 --> 00:05:02,680
Flat structure.

153
00:05:02,680 --> 00:05:04,600
No bureaucracy stacked on top of them.

154
00:05:04,600 --> 00:05:06,440
More like a trading floor than a research lab.

155
00:05:06,440 --> 00:05:08,240
People working shoulder to shoulder.

156
00:05:08,240 --> 00:05:10,080
Moving fast because there was nobody above them

157
00:05:10,080 --> 00:05:11,240
slowing the decision down.

158
00:05:11,240 --> 00:05:13,720
That detail matters more than it sounds like it should.

159
00:05:13,720 --> 00:05:16,240
Small teams building frontier class models in parallel

160
00:05:16,240 --> 00:05:18,960
only works if each team has a narrow, clear job.

161
00:05:18,960 --> 00:05:21,040
You can't have 10 people own AI broadly.

162
00:05:21,040 --> 00:05:22,800
You can have 10 people own reasoning.

163
00:05:22,800 --> 00:05:24,160
10 people own transcription.

164
00:05:24,160 --> 00:05:25,640
10 people own image generation.

165
00:05:25,640 --> 00:05:28,280
The org structure itself is already a division of labor

166
00:05:28,280 --> 00:05:30,200
before a single line of code gets written.

167
00:05:30,200 --> 00:05:32,560
And that's the real signal buried in this announcement.

168
00:05:32,560 --> 00:05:33,640
Speed like this.

169
00:05:33,640 --> 00:05:36,480
Across seven models plus a whole separate 5/4 track

170
00:05:36,480 --> 00:05:37,920
doesn't happen when everyone's building

171
00:05:37,920 --> 00:05:40,800
slightly different versions of the same general purpose thing.

172
00:05:40,800 --> 00:05:42,600
It happens when the jobs are already split.

173
00:05:42,600 --> 00:05:44,880
Small teams move fast because they're not fighting

174
00:05:44,880 --> 00:05:45,920
over what the model should be.

175
00:05:45,920 --> 00:05:46,880
They already know.

176
00:05:46,880 --> 00:05:49,040
Which means the reason and runtime split we just laid out

177
00:05:49,040 --> 00:05:50,960
wasn't an accident that emerged later.

178
00:05:50,960 --> 00:05:53,080
It was baked into how Microsoft built these models

179
00:05:53,080 --> 00:05:55,920
from day one to different design philosophies.

180
00:05:55,920 --> 00:05:58,400
So now that you know why these teams moved fast,

181
00:05:58,400 --> 00:06:00,600
let's get into why they built two different things

182
00:06:00,600 --> 00:06:02,040
instead of one better thing.

183
00:06:02,040 --> 00:06:04,520
MAI-1 is built for scale and reasoning depth.

184
00:06:04,520 --> 00:06:07,120
5/4 is built for density and deployment speed.

185
00:06:07,120 --> 00:06:08,800
That's not a difference in quality.

186
00:06:08,800 --> 00:06:10,880
It's a difference in what problem each model is solving

187
00:06:10,880 --> 00:06:12,800
before it even gets asked a question.

188
00:06:12,800 --> 00:06:14,760
Here's where you need two words that sound technical

189
00:06:14,760 --> 00:06:15,840
but really aren't.

190
00:06:15,840 --> 00:06:17,160
Dense and Sparse.

191
00:06:17,160 --> 00:06:19,240
Think about a company where every single employee

192
00:06:19,240 --> 00:06:22,040
has to sit in on every single meeting no matter what it's about.

193
00:06:22,040 --> 00:06:22,920
That's a dense model.

194
00:06:22,920 --> 00:06:25,280
Every parameter fires for every request.

195
00:06:25,280 --> 00:06:27,360
Whether the question is what's two plus two

196
00:06:27,360 --> 00:06:29,880
or restructure our entire supply chain.

197
00:06:29,880 --> 00:06:30,880
Nothing gets to sit out.

198
00:06:30,880 --> 00:06:32,680
That's 5/4 approach and on purpose.

199
00:06:32,680 --> 00:06:35,640
It's a dense transformer with a small parameter count.

200
00:06:35,640 --> 00:06:37,880
Meaning there's less total headcount in the building

201
00:06:37,880 --> 00:06:39,600
but everyone who's there is working every time.

202
00:06:39,600 --> 00:06:41,160
Now picture a different company.

203
00:06:41,160 --> 00:06:42,800
Hundreds of specialists on staff.

204
00:06:42,800 --> 00:06:45,360
But for any given meeting only the three or four people

205
00:06:45,360 --> 00:06:46,960
who actually know that topic show up.

206
00:06:46,960 --> 00:06:48,360
Everyone else stays at their desk.

207
00:06:48,360 --> 00:06:50,560
That's Sparse. That's MI1.

208
00:06:50,560 --> 00:06:52,320
It's a mixture of experts model.

209
00:06:52,320 --> 00:06:55,480
Meaning there's a massive total roster of expertise sitting inside it

210
00:06:55,480 --> 00:06:58,080
but only a slice of that roster activates per token.

211
00:06:58,080 --> 00:07:01,800
Huge capacity on paper, partial activation in practice.

212
00:07:01,800 --> 00:07:04,200
That distinction is why you can't just say bigger is smarter

213
00:07:04,200 --> 00:07:05,200
and move on.

214
00:07:05,200 --> 00:07:07,400
MI1's total size is enormous

215
00:07:07,400 --> 00:07:09,320
but it's not paying the full cost of that size

216
00:07:09,320 --> 00:07:10,600
on every single request

217
00:07:10,600 --> 00:07:13,120
because most of the model is sitting quiet at any given moment.

218
00:07:13,120 --> 00:07:15,560
If 5/4 doesn't have that luxury and doesn't need it,

219
00:07:15,560 --> 00:07:18,800
it's small enough that using all of itself every time is still cheap.

220
00:07:18,800 --> 00:07:21,600
High quality per parameter is the whole design goal.

221
00:07:21,600 --> 00:07:23,280
Nothing wasted, nothing held in reserve

222
00:07:23,280 --> 00:07:25,400
because there's nothing to spare in the first place.

223
00:07:25,400 --> 00:07:28,720
Here is why this isn't just an engineering detail you can skip past.

224
00:07:28,720 --> 00:07:31,840
Picking the wrong one for a workload doesn't just make things slightly worse.

225
00:07:31,840 --> 00:07:34,320
It actively burns money or staff's capability

226
00:07:34,320 --> 00:07:36,000
depending on which way you get it wrong.

227
00:07:36,000 --> 00:07:38,520
Send a two sentence customer question through a Sparse,

228
00:07:38,520 --> 00:07:40,320
Frontier Scale Reasoning System

229
00:07:40,320 --> 00:07:42,000
and you're paying for a roster of specialists

230
00:07:42,000 --> 00:07:44,520
to show up for a meeting that needed one intern.

231
00:07:44,520 --> 00:07:47,360
Send a genuinely complex multi-step planning problem

232
00:07:47,360 --> 00:07:48,840
through a small dense model

233
00:07:48,840 --> 00:07:51,200
and you're asking a lean, fast team

234
00:07:51,200 --> 00:07:54,480
to solve something that actually needed deep, specialized reasoning

235
00:07:54,480 --> 00:07:56,400
they don't have on staff.

236
00:07:56,400 --> 00:07:58,880
Neither failure shows up as the model was bad.

237
00:07:58,880 --> 00:08:01,200
It shows up as a budget line that's too high

238
00:08:01,200 --> 00:08:04,160
or an output that's quietly, consistently shallow.

239
00:08:04,160 --> 00:08:06,520
That's the stakes, not which model wins a benchmark

240
00:08:06,520 --> 00:08:08,240
but which one you handed the wrong job.

241
00:08:08,240 --> 00:08:11,160
And this is where the abstract idea of dense versus Sparse

242
00:08:11,160 --> 00:08:13,400
turns into something you can actually put a number on

243
00:08:13,400 --> 00:08:16,120
because once you look at what these architectures cost to run

244
00:08:16,120 --> 00:08:18,240
and what they can actually hold in memory at once,

245
00:08:18,240 --> 00:08:21,480
the gap between reason and runtime stops being conceptual

246
00:08:21,480 --> 00:08:23,880
and starts being a line item.

247
00:08:23,880 --> 00:08:25,960
The compute economics nobody talks about,

248
00:08:25,960 --> 00:08:27,440
let's put actual numbers on this.

249
00:08:27,440 --> 00:08:29,800
The story of dense versus Sparse models

250
00:08:29,800 --> 00:08:32,560
only becomes real once you see what it costs to run them.

251
00:08:32,560 --> 00:08:33,880
Start with the context window.

252
00:08:33,880 --> 00:08:37,280
MI1 reportedly holds two million tokens in a single window

253
00:08:37,280 --> 00:08:39,840
which is roughly double what GPT-5 carries.

254
00:08:39,840 --> 00:08:41,800
But this isn't just about having more room.

255
00:08:41,800 --> 00:08:44,360
It is the difference between feeding a model a single chapter

256
00:08:44,360 --> 00:08:46,560
and feeding it the entire book, the footnotes

257
00:08:46,560 --> 00:08:49,040
and the author's previous three drafts all at once.

258
00:08:49,040 --> 00:08:50,920
That window matters for the exact jobs

259
00:08:50,920 --> 00:08:52,720
MI1 is supposed to handle.

260
00:08:52,720 --> 00:08:55,760
We are talking about sprawling multi-step reasoning problems.

261
00:08:55,760 --> 00:08:57,640
We're losing context halfway through means

262
00:08:57,640 --> 00:08:59,080
the entire answer falls apart.

263
00:08:59,080 --> 00:09:00,200
Now look at the price.

264
00:09:00,200 --> 00:09:04,680
MI1 runs around $0.001 per 1,000 tokens.

265
00:09:04,680 --> 00:09:07,480
5.4 has a marginal cost that is close to zero

266
00:09:07,480 --> 00:09:09,600
once it is sitting on your local hardware.

267
00:09:09,600 --> 00:09:12,840
It is not literally zero because electricity costs money

268
00:09:12,840 --> 00:09:14,280
but it is close enough that it stops

269
00:09:14,280 --> 00:09:15,840
feeling like a meter of expense.

270
00:09:15,840 --> 00:09:17,120
That is a massive gap.

271
00:09:17,120 --> 00:09:19,920
It is the difference between paying a monthly utility bill

272
00:09:19,920 --> 00:09:22,600
and using a subscription you already bought and paid for.

273
00:09:22,600 --> 00:09:25,040
And here is why that gap exists in the first place.

274
00:09:25,040 --> 00:09:26,560
MI1 is a cloud-hosted model

275
00:09:26,560 --> 00:09:29,200
so every single request has to travel to a Dytacenter

276
00:09:29,200 --> 00:09:31,640
and hit Microsoft's infrastructure to get billed.

277
00:09:31,640 --> 00:09:34,280
Even with Sparse activation keeping the costs down

278
00:09:34,280 --> 00:09:35,840
you are still paying for the trip.

279
00:09:35,840 --> 00:09:37,320
5.4 does not make that trip at all

280
00:09:37,320 --> 00:09:40,200
once it is loaded onto a laptop or a Windows endpoint.

281
00:09:40,200 --> 00:09:41,320
You already own the hardware

282
00:09:41,320 --> 00:09:43,240
and the model is just sitting there ready to work.

283
00:09:43,240 --> 00:09:44,960
The throughput number is back this up too.

284
00:09:44,960 --> 00:09:48,840
MI1 reportedly pushes around 145 tokens per second

285
00:09:48,840 --> 00:09:51,680
on Microsoft's own Maya 200 chips.

286
00:09:51,680 --> 00:09:53,040
That is a real performance number

287
00:09:53,040 --> 00:09:54,600
rather than a marketing figure.

288
00:09:54,600 --> 00:09:56,480
And it determines if a frontier class model

289
00:09:56,480 --> 00:09:58,160
is actually usable at scale.

290
00:09:58,160 --> 00:10:00,640
It is fast enough to serve real traffic in production

291
00:10:00,640 --> 00:10:03,040
instead of just looking good during a keynote demo.

292
00:10:03,040 --> 00:10:04,600
The reason that throughput exists

293
00:10:04,600 --> 00:10:06,800
traces back to how the model was trained.

294
00:10:06,800 --> 00:10:09,680
Microsoft claims they cut training costs by 40%

295
00:10:09,680 --> 00:10:11,800
by running on Maya co-designed infrastructure

296
00:10:11,800 --> 00:10:13,640
instead of standard Nvidia clusters

297
00:10:13,640 --> 00:10:14,960
that is not just a footnote.

298
00:10:14,960 --> 00:10:16,360
Training a frontier scale model

299
00:10:16,360 --> 00:10:18,680
is one of the largest expenses in this industry

300
00:10:18,680 --> 00:10:20,520
and saving 40% changes

301
00:10:20,520 --> 00:10:22,640
what Microsoft can afford to build next.

302
00:10:22,640 --> 00:10:25,320
It also changes how aggressively they can price the models

303
00:10:25,320 --> 00:10:26,760
they have already built.

304
00:10:26,760 --> 00:10:28,320
None of this is really a technical story.

305
00:10:28,320 --> 00:10:30,800
It is a business story wearing technical clothes.

306
00:10:30,800 --> 00:10:33,000
Cost per token is not an abstract metric

307
00:10:33,000 --> 00:10:34,760
that only engineers care about.

308
00:10:34,760 --> 00:10:37,600
It is the number that decides if a project is even viable.

309
00:10:37,600 --> 00:10:39,200
If you drop the cost far enough,

310
00:10:39,200 --> 00:10:41,760
a use case that used to be too expensive to automate

311
00:10:41,760 --> 00:10:44,120
suddenly gets greenlit without a second thought.

312
00:10:44,120 --> 00:10:47,160
If the cost stays high, even a useful feature gets shelved

313
00:10:47,160 --> 00:10:48,600
because the math does not work.

314
00:10:48,600 --> 00:10:51,360
Microsoft is not just cutting costs for the sake of it.

315
00:10:51,360 --> 00:10:53,040
They are expanding the list of problems

316
00:10:53,040 --> 00:10:55,040
that are actually worth solving with AI.

317
00:10:55,040 --> 00:10:57,200
But here is the catch with these efficiency numbers.

318
00:10:57,200 --> 00:10:59,440
A low cost per token and fast throughput

319
00:10:59,440 --> 00:11:01,240
do not mean anything on their own.

320
00:11:01,240 --> 00:11:03,600
They only matter once you know where each model is supposed

321
00:11:03,600 --> 00:11:05,880
to live and what kind of request is supposed to reach it.

322
00:11:05,880 --> 00:11:07,280
Cheap and fast is great,

323
00:11:07,280 --> 00:11:08,720
but cheap and fast in the wrong place

324
00:11:08,720 --> 00:11:11,120
is just a different way to waste your money.

325
00:11:11,120 --> 00:11:13,800
Five four as the runtime, what that actually means.

326
00:11:13,800 --> 00:11:15,440
Let's define the word runtime plainly

327
00:11:15,440 --> 00:11:17,640
because it is doing a lot of work in this episode.

328
00:11:17,640 --> 00:11:19,840
A runtime is the thing that executes.

329
00:11:19,840 --> 00:11:22,080
It is not the strategist deciding what should happen,

330
00:11:22,080 --> 00:11:24,080
but the part that actually does the work right

331
00:11:24,080 --> 00:11:25,000
where the action is.

332
00:11:25,000 --> 00:11:26,360
It happens instantly.

333
00:11:26,360 --> 00:11:28,400
There is no queue, no roundtrip,

334
00:11:28,400 --> 00:11:30,880
and no waiting on a data center somewhere to weigh in.

335
00:11:30,880 --> 00:11:33,280
That is the specific job fee four was built for.

336
00:11:33,280 --> 00:11:36,320
Look at the actual sizes and the picture gets concrete fast.

337
00:11:36,320 --> 00:11:39,080
Five four mini sits at 3.8 billion parameters,

338
00:11:39,080 --> 00:11:41,760
while five four multi-model sits at 5.6 billion.

339
00:11:41,760 --> 00:11:43,800
Both of them ship under the MIT license.

340
00:11:43,800 --> 00:11:45,600
That licensing detail sounds boring

341
00:11:45,600 --> 00:11:47,920
until you are the person who has to get a model approved

342
00:11:47,920 --> 00:11:48,800
for production.

343
00:11:48,800 --> 00:11:51,240
Most frontier models come wrapped in usage restrictions

344
00:11:51,240 --> 00:11:53,480
or revenue thresholds that need a lawyer sign off

345
00:11:53,480 --> 00:11:54,640
before anyone touches them.

346
00:11:54,640 --> 00:11:56,640
MIT licensing skips that entire process

347
00:11:56,640 --> 00:11:57,720
that there is no legal friction

348
00:11:57,720 --> 00:11:59,440
and no need to check with procurement

349
00:11:59,440 --> 00:12:01,040
before you embed the model.

350
00:12:01,040 --> 00:12:04,000
You can drop five four into a product or an internal tool

351
00:12:04,000 --> 00:12:07,040
without waiting on a contract review for a runtime layer.

352
00:12:07,040 --> 00:12:09,760
That matters because these tools are supposed to be everywhere,

353
00:12:09,760 --> 00:12:12,120
rather than gated behind a vendor agreement.

354
00:12:12,120 --> 00:12:14,840
Function calling is built directly into five four mini,

355
00:12:14,840 --> 00:12:16,760
and that is what makes it function as a runtime

356
00:12:16,760 --> 00:12:18,360
instead of just a small chatbot.

357
00:12:18,360 --> 00:12:20,200
The model is not just generating text

358
00:12:20,200 --> 00:12:23,120
because it can recognize when a task requires a decision point.

359
00:12:23,120 --> 00:12:26,760
It then hands off the task to a specific tool or function

360
00:12:26,760 --> 00:12:27,680
to act on it.

361
00:12:27,680 --> 00:12:29,160
That is the difference between a model

362
00:12:29,160 --> 00:12:30,680
that describes what should happen

363
00:12:30,680 --> 00:12:32,800
and one that can actually trigger the action.

364
00:12:32,800 --> 00:12:35,280
Small agents that need to make a call and act quickly

365
00:12:35,280 --> 00:12:37,640
run perfectly on this kind of setup,

366
00:12:37,640 --> 00:12:40,040
picture where this actually shows up in the real world.

367
00:12:40,040 --> 00:12:42,440
You could have a laptop running a local coding assistant

368
00:12:42,440 --> 00:12:44,400
that checks a file for errors and fixes them

369
00:12:44,400 --> 00:12:46,440
without ever sending that data anywhere.

370
00:12:46,440 --> 00:12:48,320
A Windows endpoint could run a small agent

371
00:12:48,320 --> 00:12:49,840
that triages an incoming request

372
00:12:49,840 --> 00:12:52,480
and decides which internal tool should handle it next.

373
00:12:52,480 --> 00:12:54,680
There is no cloud call, no latency spike,

374
00:12:54,680 --> 00:12:57,080
and no data leaving the machine during the process.

375
00:12:57,080 --> 00:12:59,640
That is execution living close to the work.

376
00:12:59,640 --> 00:13:02,160
The response feels instant because there is no travel time

377
00:13:02,160 --> 00:13:05,240
involved and the device has everything it needs locally.

378
00:13:05,240 --> 00:13:06,920
This is also where data sovereignty stops

379
00:13:06,920 --> 00:13:09,320
being an abstract conversation about compliance.

380
00:13:09,320 --> 00:13:11,160
If the model runs on the endpoint,

381
00:13:11,160 --> 00:13:13,240
the request never leaves the endpoint.

382
00:13:13,240 --> 00:13:15,640
That is not a policy you have to enforce after the fact.

383
00:13:15,640 --> 00:13:18,160
It is just how the architecture of the system works.

384
00:13:18,160 --> 00:13:20,080
But we have to be honest about the limits here.

385
00:13:20,080 --> 00:13:22,240
Execution alone does not explain intelligence.

386
00:13:22,240 --> 00:13:24,440
A runtime can act fast and act locally,

387
00:13:24,440 --> 00:13:26,480
but it cannot plan five moves ahead

388
00:13:26,480 --> 00:13:28,720
through a genuinely complicated problem.

389
00:13:28,720 --> 00:13:30,760
Five fork and catch an error or trigger a function

390
00:13:30,760 --> 00:13:32,200
to resolve a routine case.

391
00:13:32,200 --> 00:13:34,440
It cannot sit there and reason through something

392
00:13:34,440 --> 00:13:36,720
that requires holding a massive set of trade-offs

393
00:13:36,720 --> 00:13:37,840
in its head at once.

394
00:13:37,840 --> 00:13:39,000
That is not a flaw in the system.

395
00:13:39,000 --> 00:13:39,840
It is the design.

396
00:13:39,840 --> 00:13:42,520
The deep planning work was never supposed to live at the edge.

397
00:13:42,520 --> 00:13:44,040
That is the job for MI1.

398
00:13:44,040 --> 00:13:46,000
And we need to see what actual reason looks like

399
00:13:46,000 --> 00:13:48,160
once you get past the marketing label.

400
00:13:48,160 --> 00:13:51,040
MI1 as the reason, what that actually means.

401
00:13:51,040 --> 00:13:53,480
Reasoning means something very specific in this context.

402
00:13:53,480 --> 00:13:54,640
It isn't the part of the brain

403
00:13:54,640 --> 00:13:56,400
that remembers what happened yesterday.

404
00:13:56,400 --> 00:13:58,200
It's the part that decides what should happen next.

405
00:13:58,200 --> 00:14:00,320
It weighs options, it holds trade-offs,

406
00:14:00,320 --> 00:14:03,560
and it plans out five steps before a single line of code is written.

407
00:14:03,560 --> 00:14:05,680
That is a different job than what five orders.

408
00:14:05,680 --> 00:14:07,040
And because the job is different,

409
00:14:07,040 --> 00:14:08,880
it needs a different model underneath it.

410
00:14:08,880 --> 00:14:10,840
Microsoft put that job into my thinking one.

411
00:14:10,840 --> 00:14:13,000
It runs on 35 billion active parameters

412
00:14:13,000 --> 00:14:16,520
with a context window of 256,000 tokens.

413
00:14:16,520 --> 00:14:18,480
On the AIME 2025 math benchmark,

414
00:14:18,480 --> 00:14:20,040
it scores 97%.

415
00:14:20,040 --> 00:14:22,120
That isn't just a high score on a test.

416
00:14:22,120 --> 00:14:24,640
That specific benchmark is designed to trip models up

417
00:14:24,640 --> 00:14:26,840
with layered multi-step problems.

418
00:14:26,840 --> 00:14:30,320
Hitting 97% means the model can hold a complex problem together

419
00:14:30,320 --> 00:14:33,080
across a long chain of logic without losing the thread.

420
00:14:33,080 --> 00:14:36,120
But there is a detail here that matters more than the math scores.

421
00:14:36,120 --> 00:14:39,040
MI thinking one was built with zero distillation.

422
00:14:39,040 --> 00:14:41,160
In this industry, most models are built

423
00:14:41,160 --> 00:14:43,920
by training a small system to mimic a big one.

424
00:14:43,920 --> 00:14:46,600
You essentially teach the student to copy the teacher's homework.

425
00:14:46,600 --> 00:14:49,160
That's distillation, it's fast, but it has a hidden cost.

426
00:14:49,160 --> 00:14:51,120
You inherit every bias, every gap,

427
00:14:51,120 --> 00:14:52,760
and every licensing red flag

428
00:14:52,760 --> 00:14:54,640
that was baked into the original model.

429
00:14:54,640 --> 00:14:56,600
Zero distillation means none of that.

430
00:14:56,600 --> 00:14:58,880
MI thinking one was built from the ground up,

431
00:14:58,880 --> 00:15:01,960
using data with a clean, commercially licensed lineage.

432
00:15:01,960 --> 00:15:04,080
Nothing was borrowed from a black box source.

433
00:15:04,080 --> 00:15:07,200
If you work in a regulated industry like healthcare or finance,

434
00:15:07,200 --> 00:15:08,680
this isn't just a nice feature.

435
00:15:08,680 --> 00:15:11,840
It is the difference between a model your compliance team signs off on

436
00:15:11,840 --> 00:15:13,680
and one that gets stuck in review forever

437
00:15:13,680 --> 00:15:15,840
because nobody knows where its knowledge came from.

438
00:15:15,840 --> 00:15:17,440
This isn't just about math either.

439
00:15:17,440 --> 00:15:21,560
On SWEBench Pro, which tests real world software engineering,

440
00:15:21,560 --> 00:15:23,720
MI thinking one hits 53%.

441
00:15:23,720 --> 00:15:25,800
That puts it right next to Opus class models

442
00:15:25,800 --> 00:15:29,000
for a long time, Opus has been the gold standard for complex coding.

443
00:15:29,000 --> 00:15:32,400
Landing in that range proves MI one can handle the messy, ambiguous,

444
00:15:32,400 --> 00:15:35,000
multi-file reasoning that actual engineering requires.

445
00:15:35,000 --> 00:15:36,840
But we need to break a common instinct here.

446
00:15:36,840 --> 00:15:38,520
We've talked about how 5.4 is efficient,

447
00:15:38,520 --> 00:15:40,520
but bigger doesn't always mean better.

448
00:15:40,520 --> 00:15:42,440
Sparse activation proves that smaller footprints

449
00:15:42,440 --> 00:15:44,160
can outperform brute force.

450
00:15:44,160 --> 00:15:47,360
However, deeper reasoning still requires more active compute.

451
00:15:47,360 --> 00:15:48,960
That part is not negotiable.

452
00:15:48,960 --> 00:15:51,160
You cannot compress genuine, multi-step planning

453
00:15:51,160 --> 00:15:53,960
into a few billion parameters and expect the same result.

454
00:15:53,960 --> 00:15:56,960
Those 35 billion active parameters aren't there for show.

455
00:15:56,960 --> 00:15:59,800
They are the literal cost of holding a hard problem together

456
00:15:59,800 --> 00:16:01,080
long enough to solve it.

457
00:16:01,080 --> 00:16:03,360
This is where the two layers stop being separate stories.

458
00:16:03,360 --> 00:16:06,080
5.4 executes its fast, cheap, and local.

459
00:16:06,080 --> 00:16:08,720
MI one reasons its deep, deliberate, and expensive.

460
00:16:08,720 --> 00:16:10,200
Neither one replaces the other

461
00:16:10,200 --> 00:16:11,960
and neither one is complete on its own.

462
00:16:11,960 --> 00:16:14,440
The real architecture isn't about picking one model.

463
00:16:14,440 --> 00:16:15,560
It's about the handoff.

464
00:16:15,560 --> 00:16:18,240
It's about how a request starts at the runtime layer

465
00:16:18,240 --> 00:16:21,000
and moves to the reason layer the moment it needs real judgment.

466
00:16:21,000 --> 00:16:23,040
That handoff is the part nobody is talking about.

467
00:16:23,040 --> 00:16:25,880
And it's the most important piece of engineering to understand.

468
00:16:25,880 --> 00:16:28,360
The handoff, how reasoning becomes execution.

469
00:16:28,360 --> 00:16:30,200
Here's what that handoff looks like when you treat it

470
00:16:30,200 --> 00:16:32,000
as a system instead of an abstraction.

471
00:16:32,000 --> 00:16:33,120
A request comes in.

472
00:16:33,120 --> 00:16:36,040
Something has to look at it first before either model touches it

473
00:16:36,040 --> 00:16:37,760
to decide what it actually is.

474
00:16:37,760 --> 00:16:39,280
It isn't looking at the surface level.

475
00:16:39,280 --> 00:16:41,800
It's looking at what kind of problem it represents.

476
00:16:41,800 --> 00:16:44,400
Is this a two-line question with an obvious answer?

477
00:16:44,400 --> 00:16:46,040
Or is this something that needs five steps

478
00:16:46,040 --> 00:16:48,200
of planning before a response makes sense?

479
00:16:48,200 --> 00:16:49,760
That classification happens first.

480
00:16:49,760 --> 00:16:52,320
Then and only then the request gets routed.

481
00:16:52,320 --> 00:16:54,560
Simple, high volume tasks go to 5.4.

482
00:16:54,560 --> 00:16:55,840
It might be running on your device

483
00:16:55,840 --> 00:16:58,040
or hosted as a small model in Foundry.

484
00:16:58,040 --> 00:17:01,200
Complex multi-step reasoning gets escalated to MAI1.

485
00:17:01,200 --> 00:17:02,280
The pattern is simple.

486
00:17:02,280 --> 00:17:03,480
Classify, then route.

487
00:17:03,480 --> 00:17:04,560
But here's the problem.

488
00:17:04,560 --> 00:17:07,080
Everyone wants to talk about which model is smarter.

489
00:17:07,080 --> 00:17:09,400
Almost nobody wants to talk about the routing layer.

490
00:17:09,400 --> 00:17:11,520
This is the infrastructure that decides which model

491
00:17:11,520 --> 00:17:13,320
even sees the request in the first place.

492
00:17:13,320 --> 00:17:14,600
It isn't a glamorous build.

493
00:17:14,600 --> 00:17:16,160
It won't get a keynote moment.

494
00:17:16,160 --> 00:17:18,560
But it is the difference between a system that works

495
00:17:18,560 --> 00:17:20,720
and two expensive models sitting next to each other

496
00:17:20,720 --> 00:17:22,480
with no logic connecting them.

497
00:17:22,480 --> 00:17:24,200
Think about a support ticket system.

498
00:17:24,200 --> 00:17:25,920
80% of the tickets are routine.

499
00:17:25,920 --> 00:17:28,440
People want to reset passwords, check order status,

500
00:17:28,440 --> 00:17:29,720
or find an invoice.

501
00:17:29,720 --> 00:17:31,160
These questions have one clear answer

502
00:17:31,160 --> 00:17:32,840
and don't require weighing trade-offs.

503
00:17:32,840 --> 00:17:35,360
Those tickets resolve entirely at the 5.4 layer.

504
00:17:35,360 --> 00:17:37,560
They are fast, they are cheap, and they are done.

505
00:17:37,560 --> 00:17:39,240
The other 20% are different.

506
00:17:39,240 --> 00:17:42,120
Maybe it's a billing dispute that touches three different systems

507
00:17:42,120 --> 00:17:44,840
or a technical issue that needs a root cause analysis

508
00:17:44,840 --> 00:17:46,520
across a chain of dependencies.

509
00:17:46,520 --> 00:17:48,560
Those get escalated and MAI1 picks them up.

510
00:17:48,560 --> 00:17:50,160
That 80/20 split isn't a guess.

511
00:17:50,160 --> 00:17:52,680
It matches the data we see across the industry.

512
00:17:52,680 --> 00:17:55,080
Most requests are routine and only a small minority

513
00:17:55,080 --> 00:17:56,240
need deep reasoning.

514
00:17:56,240 --> 00:17:59,000
The routing layer's entire job is telling the difference

515
00:17:59,000 --> 00:18:02,480
correctly every single time without a human checking the work.

516
00:18:02,480 --> 00:18:04,400
This matters for more than just the budget.

517
00:18:04,400 --> 00:18:07,120
Yes, it's cheaper to solve routine tickets with a small model

518
00:18:07,120 --> 00:18:09,440
instead of paying frontier rates for a simple question.

519
00:18:09,440 --> 00:18:10,960
But the real win is latency.

520
00:18:10,960 --> 00:18:13,400
Think about the user experience for that 80%.

521
00:18:13,400 --> 00:18:15,920
They ask a question and the answer comes back almost instantly

522
00:18:15,920 --> 00:18:18,160
because the request never left the local layer.

523
00:18:18,160 --> 00:18:19,080
There is no queue.

524
00:18:19,080 --> 00:18:20,800
There is no wait for a massive system

525
00:18:20,800 --> 00:18:23,360
to spin up its experts for a question that didn't need them.

526
00:18:23,360 --> 00:18:25,280
The handoff to MI1 happens invisibly.

527
00:18:25,280 --> 00:18:28,480
It only happens for the cases that actually justify the wait.

528
00:18:28,480 --> 00:18:31,200
Most users will never notice a handoff even occurred.

529
00:18:31,200 --> 00:18:33,800
They just experience a system that feels fast all the time.

530
00:18:33,800 --> 00:18:35,960
It only takes longer on the genuinely hard stuff

531
00:18:35,960 --> 00:18:37,480
and they never have to know why.

532
00:18:37,480 --> 00:18:39,160
That invisibility is the whole point.

533
00:18:39,160 --> 00:18:41,640
A well-built routing layer doesn't announce itself.

534
00:18:41,640 --> 00:18:44,200
It makes the system feel like one single intelligence

535
00:18:44,200 --> 00:18:46,320
even though two different models are doing the work.

536
00:18:46,320 --> 00:18:47,600
And here is the thing to consider.

537
00:18:47,600 --> 00:18:50,120
This pattern of classifying, routing, and escalating

538
00:18:50,120 --> 00:18:52,920
only when necessary isn't something Microsoft invented

539
00:18:52,920 --> 00:18:53,880
in a vacuum.

540
00:18:53,880 --> 00:18:55,960
It isn't unique to MI1 and 5.4.

541
00:18:55,960 --> 00:18:57,240
This pattern is showing up everywhere.

542
00:18:57,240 --> 00:18:59,320
Company after company is moving toward this,

543
00:18:59,320 --> 00:19:01,280
regardless of which models they use.

544
00:19:01,280 --> 00:19:02,560
And that raises a real question.

545
00:19:02,560 --> 00:19:04,800
If everyone is converging on the same architecture,

546
00:19:04,800 --> 00:19:07,320
what does that tell you about where this is headed?

547
00:19:07,320 --> 00:19:09,560
Why the industry already agrees on this split?

548
00:19:09,560 --> 00:19:12,360
This shift is bigger than one company's product strategy.

549
00:19:12,360 --> 00:19:15,080
Gardner is projecting that by 2027,

550
00:19:15,080 --> 00:19:18,840
organizations will use small, task-specific models three times

551
00:19:18,840 --> 00:19:21,960
more often than general-purpose LLMs, not slightly more.

552
00:19:21,960 --> 00:19:22,880
Three times.

553
00:19:22,880 --> 00:19:25,200
That isn't a niche trend tucked into a research footnote.

554
00:19:25,200 --> 00:19:27,720
It's analysts describing what's about to become the default

555
00:19:27,720 --> 00:19:29,520
shape of enterprise AI.

556
00:19:29,520 --> 00:19:32,200
The do everything models, stop being the norm.

557
00:19:32,200 --> 00:19:34,040
The task-specific small models take over

558
00:19:34,040 --> 00:19:35,680
the majority of the workload.

559
00:19:35,680 --> 00:19:37,120
And this isn't some future prediction

560
00:19:37,120 --> 00:19:38,680
with nothing behind it yet.

561
00:19:38,680 --> 00:19:41,400
In AI-major markets, 68% of enterprises

562
00:19:41,400 --> 00:19:43,720
are already running at least one small model

563
00:19:43,720 --> 00:19:44,640
in production today.

564
00:19:44,640 --> 00:19:46,560
Not piloting, not testing in a sandbox,

565
00:19:46,560 --> 00:19:49,240
running it in production right now.

566
00:19:49,240 --> 00:19:51,000
That number alone should tell you the industry

567
00:19:51,000 --> 00:19:53,600
didn't wait around for a keynote to figure this out.

568
00:19:53,600 --> 00:19:56,280
Companies were already building toward exactly this split

569
00:19:56,280 --> 00:20:00,000
before Microsoft ever said the words MI1 or 5-4 out loud.

570
00:20:00,000 --> 00:20:03,120
Here's the economic driver underneath all of it, stated plainly.

571
00:20:03,120 --> 00:20:05,040
Serving a 7-billion parameter model

572
00:20:05,040 --> 00:20:07,240
can be 10 to 30 times cheaper than serving a model

573
00:20:07,240 --> 00:20:10,400
in the 70-billion plus range, 10 to 30 times.

574
00:20:10,400 --> 00:20:12,880
That isn't a rounding error you absorb into overhead.

575
00:20:12,880 --> 00:20:15,560
It's the difference between a workload that scales sustainably

576
00:20:15,560 --> 00:20:18,960
and one that gets quietly killed in a budget review six months in.

577
00:20:18,960 --> 00:20:20,840
Because the ongoing cost never made sense

578
00:20:20,840 --> 00:20:22,120
once volume actually showed up.

579
00:20:22,120 --> 00:20:24,360
Once you see that math, the industry wide shift

580
00:20:24,360 --> 00:20:25,760
stops looking like a trend.

581
00:20:25,760 --> 00:20:28,880
It starts looking like the only rational move available.

582
00:20:28,880 --> 00:20:31,560
If a small model handles the routine 80% of requests

583
00:20:31,560 --> 00:20:33,000
that are fraction of the cost,

584
00:20:33,000 --> 00:20:34,680
and the frontier model only gets pulled in

585
00:20:34,680 --> 00:20:36,480
for the genuinely hard 20%,

586
00:20:36,480 --> 00:20:38,520
you aren't choosing between quality and savings,

587
00:20:38,520 --> 00:20:39,520
you're getting both.

588
00:20:39,520 --> 00:20:42,200
Because you stopped paying frontier prices for questions

589
00:20:42,200 --> 00:20:44,640
that never needed frontier reasoning in the first place.

590
00:20:44,640 --> 00:20:46,400
So here's the point worth being honest about.

591
00:20:46,400 --> 00:20:48,080
Microsoft didn't invent this pattern.

592
00:20:48,080 --> 00:20:49,560
This isn't some proprietary insight

593
00:20:49,560 --> 00:20:51,080
that only came out of a Redmond Lab.

594
00:20:51,080 --> 00:20:54,200
What Microsoft did was productize it at their own scale,

595
00:20:54,200 --> 00:20:55,560
running on their own silicon.

596
00:20:55,560 --> 00:20:59,040
Maya chips underneath MIA1, MIT licensed 5-4 models

597
00:20:59,040 --> 00:21:00,320
built to sit on endpoints,

598
00:21:00,320 --> 00:21:02,040
and the whole reason and runtime split

599
00:21:02,040 --> 00:21:04,080
dressed up in shipped as a coordinated platform.

600
00:21:04,080 --> 00:21:06,440
But the underlying logic, small models for volume,

601
00:21:06,440 --> 00:21:07,760
large models for depth,

602
00:21:07,760 --> 00:21:10,280
was already the direction the entire industry was moving.

603
00:21:10,280 --> 00:21:11,920
Microsoft just built the most visible,

604
00:21:11,920 --> 00:21:14,040
most vertically integrated version of it,

605
00:21:14,040 --> 00:21:15,480
which raises an uncomfortable question

606
00:21:15,480 --> 00:21:17,240
for anyone who hasn't caught up yet.

607
00:21:17,240 --> 00:21:20,000
If this split is already the industry consensus,

608
00:21:20,000 --> 00:21:22,560
and the data shows 68% adoption

609
00:21:22,560 --> 00:21:25,400
with a three-fold shift coming within a couple of years,

610
00:21:25,400 --> 00:21:27,080
what does it actually cost a company

611
00:21:27,080 --> 00:21:29,360
that hasn't built this architecture on purpose?

612
00:21:29,360 --> 00:21:31,880
What happens when every request, simple or complex,

613
00:21:31,880 --> 00:21:34,640
still gets rooted through the exact same frontier model?

614
00:21:34,640 --> 00:21:36,440
Because nobody ever built the routing logic

615
00:21:36,440 --> 00:21:37,600
to do anything else?

616
00:21:37,600 --> 00:21:38,840
That isn't a hypothetical.

617
00:21:38,840 --> 00:21:40,840
That's the default state most organizations

618
00:21:40,840 --> 00:21:42,080
are still sitting in right now.

619
00:21:42,080 --> 00:21:46,120
And it's worth looking at exactly what that default costs.

620
00:21:46,120 --> 00:21:48,280
The old model, one model to rule everything,

621
00:21:48,280 --> 00:21:49,960
here's what that default actually looks like

622
00:21:49,960 --> 00:21:52,040
inside a company that never built the split,

623
00:21:52,040 --> 00:21:53,520
every request comes in,

624
00:21:53,520 --> 00:21:55,360
and every request goes to the same place,

625
00:21:55,360 --> 00:21:58,520
a frontier model, general purpose, expensive.

626
00:21:58,520 --> 00:22:00,800
Handling a two-sentence password reset

627
00:22:00,800 --> 00:22:03,360
the exact same way it handles a genuinely hard

628
00:22:03,360 --> 00:22:04,720
multi-step planning problem.

629
00:22:04,720 --> 00:22:07,400
There's no classification step, no routing logic,

630
00:22:07,400 --> 00:22:09,440
no decision about what a request actually needs.

631
00:22:09,440 --> 00:22:11,880
There's just one door, and everything walks through it.

632
00:22:11,880 --> 00:22:13,440
This is the floor default,

633
00:22:13,440 --> 00:22:16,160
and it's still the reality for most organizations right now,

634
00:22:16,160 --> 00:22:18,080
not because anyone chose it deliberately,

635
00:22:18,080 --> 00:22:19,920
but because nobody built anything else.

636
00:22:19,920 --> 00:22:21,800
When you only have one model available,

637
00:22:21,800 --> 00:22:24,280
route everything through it isn't a strategy.

638
00:22:24,280 --> 00:22:26,040
It's just what happens by omission.

639
00:22:26,040 --> 00:22:27,760
Here's what that omission costs.

640
00:22:27,760 --> 00:22:29,320
Companies stuck on this pattern report

641
00:22:29,320 --> 00:22:33,720
monthly cloud AI bills running $50,000 to $100,000 or more.

642
00:22:33,720 --> 00:22:36,160
For workloads that never needed frontier level reasoning

643
00:22:36,160 --> 00:22:37,200
in the first place.

644
00:22:37,200 --> 00:22:39,720
That isn't the cost of doing hard, valuable work.

645
00:22:39,720 --> 00:22:41,680
It's the cost of asking an expensive specialist

646
00:22:41,680 --> 00:22:43,960
to sit through a meeting that only needed an intern.

647
00:22:43,960 --> 00:22:45,760
Over and over, thousands of times a month

648
00:22:45,760 --> 00:22:47,640
because there was no cheaper option on the roster,

649
00:22:47,640 --> 00:22:49,560
and the cost isn't only financial.

650
00:22:49,560 --> 00:22:51,640
There's a latency tax buried in here too,

651
00:22:51,640 --> 00:22:53,440
one that's easy to miss because it doesn't show up

652
00:22:53,440 --> 00:22:54,400
on an invoice.

653
00:22:54,400 --> 00:22:56,600
Picture a model built to hold a two million token

654
00:22:56,600 --> 00:22:58,720
context window, architected to reason

655
00:22:58,720 --> 00:23:01,520
across entire code bases and sprawling document sets,

656
00:23:01,520 --> 00:23:03,840
and then someone asks it a two-sentence question.

657
00:23:03,840 --> 00:23:05,240
The system still has to spin up,

658
00:23:05,240 --> 00:23:07,400
it still has to prepare for the possibility of holding

659
00:23:07,400 --> 00:23:09,920
that much context, even though this particular request

660
00:23:09,920 --> 00:23:11,200
needed almost none of it.

661
00:23:11,200 --> 00:23:12,400
The user waits on infrastructure

662
00:23:12,400 --> 00:23:15,200
that was built for a different kind of problem entirely.

663
00:23:15,200 --> 00:23:16,840
Here's the belief worth breaking,

664
00:23:16,840 --> 00:23:19,040
and it's an uncomfortable one for a lot of teams.

665
00:23:19,040 --> 00:23:21,040
When the setup feels slow or expensive,

666
00:23:21,040 --> 00:23:22,600
the instinct is to blame the model.

667
00:23:22,600 --> 00:23:24,760
Upgraded, tune it, throw more compute at it,

668
00:23:24,760 --> 00:23:26,480
but that's misdiagnosing the problem.

669
00:23:26,480 --> 00:23:29,480
This isn't a technology failure, it's an architecture failure.

670
00:23:29,480 --> 00:23:30,960
The model was never the bottleneck.

671
00:23:30,960 --> 00:23:32,680
The bottleneck was the absence of a system

672
00:23:32,680 --> 00:23:34,120
that knew which request deserved

673
00:23:34,120 --> 00:23:36,800
that model's full capability and which ones didn't.

674
00:23:36,800 --> 00:23:38,960
A frontier model performing exactly as designed

675
00:23:38,960 --> 00:23:41,560
on a task that never needed frontier level reasoning

676
00:23:41,560 --> 00:23:42,520
isn't a broken model.

677
00:23:42,520 --> 00:23:45,240
It's a broken decision about where that model belongs.

678
00:23:45,240 --> 00:23:47,640
No amount of upgrading fixes a routing problem

679
00:23:47,640 --> 00:23:49,800
because the model was never what needed fixing.

680
00:23:49,800 --> 00:23:52,960
This is exactly the Gap MI1 and 5.4 were built to close.

681
00:23:52,960 --> 00:23:55,040
Not by making one model do everything better,

682
00:23:55,040 --> 00:23:57,120
but by making sure everything doesn't have to go through

683
00:23:57,120 --> 00:23:58,760
one model in the first place.

684
00:23:58,760 --> 00:24:01,720
To understand how that Gap actually gets closed in practice,

685
00:24:01,720 --> 00:24:04,400
it helps to look inside what each of these two models

686
00:24:04,400 --> 00:24:07,080
is actually built from, starting with the one designed

687
00:24:07,080 --> 00:24:09,160
to live closest to the work.

688
00:24:09,160 --> 00:24:11,400
Inside 5.4's actual architecture.

689
00:24:11,400 --> 00:24:14,560
Let's open this up and look at what is sitting inside 5.4

690
00:24:14,560 --> 00:24:16,680
because the architecture explains why it can live

691
00:24:16,680 --> 00:24:18,520
at the edge in the first place.

692
00:24:18,520 --> 00:24:20,720
It starts with grouped query attention.

693
00:24:20,720 --> 00:24:23,120
In a normal transformer, every single query head

694
00:24:23,120 --> 00:24:25,120
does its own separate lookup work.

695
00:24:25,120 --> 00:24:28,160
It is thorough, but it is also expensive for your memory.

696
00:24:28,160 --> 00:24:29,800
Grouped query attention changes that

697
00:24:29,800 --> 00:24:32,800
by having clusters of query heads share the same lookup work

698
00:24:32,800 --> 00:24:34,520
instead of each one duplicating it.

699
00:24:34,520 --> 00:24:37,760
The same basic job gets done, but you spend less memory doing it.

700
00:24:37,760 --> 00:24:39,920
This is not a shortcut that hurts quality.

701
00:24:39,920 --> 00:24:41,760
It is a shortcut that removes waste

702
00:24:41,760 --> 00:24:43,360
that was never buying you anything.

703
00:24:43,360 --> 00:24:45,680
When you pair that with a 200,000 word vocabulary

704
00:24:45,680 --> 00:24:48,240
and shared input output embeddings, the model changes.

705
00:24:48,240 --> 00:24:50,400
A bigger vocabulary means the model can represent

706
00:24:50,400 --> 00:24:52,320
language more precisely without breaking words

707
00:24:52,320 --> 00:24:53,720
into awkward fragments.

708
00:24:53,720 --> 00:24:55,760
By sharing the embeddings between input and output,

709
00:24:55,760 --> 00:24:57,840
the model does not have to store two separate copies

710
00:24:57,840 --> 00:25:00,760
of the same vocabulary knowledge for reading and writing.

711
00:25:00,760 --> 00:25:02,800
It reuses the same table for both directions.

712
00:25:02,800 --> 00:25:05,440
This means fewer parameters are spent on redundancy

713
00:25:05,440 --> 00:25:06,880
and more of the model's budget goes

714
00:25:06,880 --> 00:25:09,000
toward actually understanding what you are asking.

715
00:25:09,000 --> 00:25:11,520
The multi-modal version follows a similar logic.

716
00:25:11,520 --> 00:25:13,760
5.4 multi-modal is not just a text model

717
00:25:13,760 --> 00:25:16,120
with a camera bolted on as an afterthought.

718
00:25:16,120 --> 00:25:19,000
It runs a vision encoder and an audio encoder side by side

719
00:25:19,000 --> 00:25:21,440
and both feed into that same compact backbone

720
00:25:21,440 --> 00:25:23,800
using a mixture of low-res approach.

721
00:25:23,800 --> 00:25:25,720
Instead of training three entirely separate models

722
00:25:25,720 --> 00:25:28,800
for text, image and audio, Microsoft trains small,

723
00:25:28,800 --> 00:25:30,920
lightweight adapters that plug into the core model

724
00:25:30,920 --> 00:25:32,920
depending on what kind of input shows up.

725
00:25:32,920 --> 00:25:35,240
You have one backbone and three sets of ears

726
00:25:35,240 --> 00:25:37,920
and each one only activates when it is actually needed.

727
00:25:37,920 --> 00:25:40,320
There is one detail here that is genuinely surprising.

728
00:25:40,320 --> 00:25:42,480
5.4 was trained on five trillion tokens.

729
00:25:42,480 --> 00:25:44,440
That number sounds huge until you compare it

730
00:25:44,440 --> 00:25:46,840
to frontier scale models that often train

731
00:25:46,840 --> 00:25:48,520
on two or three times that amount.

732
00:25:48,520 --> 00:25:51,520
5.4 is working with a noticeably smaller training diet

733
00:25:51,520 --> 00:25:53,160
than the giants it competes with.

734
00:25:53,160 --> 00:25:54,680
So why does it still perform well?

735
00:25:54,680 --> 00:25:56,560
It works because those five trillion tokens

736
00:25:56,560 --> 00:25:58,680
are heavily synthetic and heavily curated.

737
00:25:58,680 --> 00:26:01,120
Microsoft is not just scraping the open internet

738
00:26:01,120 --> 00:26:03,160
and hoping the quality averages out.

739
00:26:03,160 --> 00:26:05,080
They are generating targeted training data

740
00:26:05,080 --> 00:26:07,080
specifically designed to teach the model

741
00:26:07,080 --> 00:26:09,040
math coding and reasoning patterns.

742
00:26:09,040 --> 00:26:11,640
They filter hard for quality before any of it gets used.

743
00:26:11,640 --> 00:26:13,200
This is quality of a volume.

744
00:26:13,200 --> 00:26:15,680
A frontier model might need 10 times the raw data

745
00:26:15,680 --> 00:26:17,320
because a huge chunk of what it eats

746
00:26:17,320 --> 00:26:19,040
is noisy or repetitive.

747
00:26:19,040 --> 00:26:21,240
5.4's approach front loads the curation work

748
00:26:21,240 --> 00:26:23,680
so every token it sees is pulling more weight.

749
00:26:23,680 --> 00:26:25,560
You have less data but that data was built

750
00:26:25,560 --> 00:26:28,160
to teach exactly what the model needed to learn.

751
00:26:28,160 --> 00:26:30,160
Then there is SambaY.

752
00:26:30,160 --> 00:26:32,320
This variant is a genuine architectural departure

753
00:26:32,320 --> 00:26:34,080
rather than just a size adjustment.

754
00:26:34,080 --> 00:26:35,960
SambaY is a hybrid design that blends

755
00:26:35,960 --> 00:26:39,160
Mamba style state space layers in with traditional attention.

756
00:26:39,160 --> 00:26:41,640
State space layers process long sequences differently

757
00:26:41,640 --> 00:26:44,120
than attention does and they are often more efficient

758
00:26:44,120 --> 00:26:45,440
for certain kinds of context.

759
00:26:45,440 --> 00:26:48,000
By combining the two, Microsoft is testing

760
00:26:48,000 --> 00:26:49,960
whether the next efficiency gain comes

761
00:26:49,960 --> 00:26:51,760
from a smarter core architecture

762
00:26:51,760 --> 00:26:54,240
instead of just smarter training data.

763
00:26:54,240 --> 00:26:55,800
When you put all of this together,

764
00:26:55,800 --> 00:26:58,560
you get a model built to run at the edge

765
00:26:58,560 --> 00:27:00,720
on modest hardware without the cloud.

766
00:27:00,720 --> 00:27:03,080
But a runtime that lives at the edge only matters

767
00:27:03,080 --> 00:27:04,840
if it can connect to something bigger

768
00:27:04,840 --> 00:27:06,200
when a request outgrows it.

769
00:27:06,200 --> 00:27:09,000
That is where MI1's architecture comes into the picture.

770
00:27:09,000 --> 00:27:11,160
Inside MI1's actual architecture.

771
00:27:11,160 --> 00:27:13,320
Now we flip to the other side of the stack.

772
00:27:13,320 --> 00:27:15,920
MI1's architecture is solving a completely different problem

773
00:27:15,920 --> 00:27:17,520
than the one 5.4 just solved.

774
00:27:17,520 --> 00:27:19,720
It uses a sparse mixture of expert system.

775
00:27:19,720 --> 00:27:21,560
Imagine a roster of specialists.

776
00:27:21,560 --> 00:27:24,040
Way more than you would ever need in one room at once.

777
00:27:24,040 --> 00:27:26,520
Where each one is trained on a different slice of knowledge.

778
00:27:26,520 --> 00:27:28,600
When a request comes in, the system does not wake up

779
00:27:28,600 --> 00:27:29,600
the whole roster.

780
00:27:29,600 --> 00:27:31,680
It picks a small handful of experts relevant

781
00:27:31,680 --> 00:27:35,000
to that specific question and only those experts do any work.

782
00:27:35,000 --> 00:27:37,080
Everyone else on the roster stays idle.

783
00:27:37,080 --> 00:27:39,800
That is the many experts, few active structure.

784
00:27:39,800 --> 00:27:42,280
You have massive total capacity sitting on the shelf

785
00:27:42,280 --> 00:27:45,120
but only a sliver of it is switched on for any single token.

786
00:27:45,120 --> 00:27:46,920
This is the key to controlling cost.

787
00:27:46,920 --> 00:27:49,240
A dense model of a similar size would have to activate

788
00:27:49,240 --> 00:27:51,200
everything for every single request.

789
00:27:51,200 --> 00:27:53,040
That is fine when the total size is small

790
00:27:53,040 --> 00:27:55,480
but it stops being fine once the model gets frontier large.

791
00:27:55,480 --> 00:27:57,480
At that scale you would be paying full compute

792
00:27:57,480 --> 00:27:59,720
for the entire roster on every query

793
00:27:59,720 --> 00:28:02,800
whether the question needed three specialists or 300.

794
00:28:02,800 --> 00:28:05,160
Sparse activation breaks that link.

795
00:28:05,160 --> 00:28:07,680
MI1 can carry an enormous total parameter count

796
00:28:07,680 --> 00:28:10,520
without paying full price for that capacity on every request.

797
00:28:10,520 --> 00:28:13,480
You get the scale without inheriting the normal price tag.

798
00:28:13,480 --> 00:28:15,240
Then you have to layer in the hardware side.

799
00:28:15,240 --> 00:28:18,680
Microsoft co-designed MI1 with its own Maya 200 chips

800
00:28:18,680 --> 00:28:21,600
and the reported payoff is a 1.4 times gain in performance

801
00:28:21,600 --> 00:28:23,520
per what compared to generic GPUs.

802
00:28:23,520 --> 00:28:25,200
This is not just a small tuning win.

803
00:28:25,200 --> 00:28:27,960
It is the difference between squeezing gains out of hardware

804
00:28:27,960 --> 00:28:29,640
you did not build for this job

805
00:28:29,640 --> 00:28:31,480
and using hardware that was shaped around

806
00:28:31,480 --> 00:28:34,200
the model's actual activation pattern from the start.

807
00:28:34,200 --> 00:28:36,120
This co-design is a strategic move.

808
00:28:36,120 --> 00:28:37,400
Not just an engineering flex,

809
00:28:37,400 --> 00:28:38,920
anyone can rent GPU capacity

810
00:28:38,920 --> 00:28:41,680
because that is a commodity available to whoever has the budget.

811
00:28:41,680 --> 00:28:43,480
But designing your chip and your model together

812
00:28:43,480 --> 00:28:46,360
so the activation pattern maps onto the silicon underneath it

813
00:28:46,360 --> 00:28:48,640
is not something a competitor can just buy.

814
00:28:48,640 --> 00:28:51,080
It is a mode built out of years of coordinated engineering

815
00:28:51,080 --> 00:28:53,640
between teams that do not normally sit in the same building.

816
00:28:53,640 --> 00:28:55,640
Copying the model architecture is possible

817
00:28:55,640 --> 00:28:58,040
but copying the chip model relationship behind it

818
00:28:58,040 --> 00:28:59,240
is a much longer road.

819
00:28:59,240 --> 00:29:02,320
And that is exactly the point Suleiman has been making.

820
00:29:02,320 --> 00:29:04,000
The goal is not just building good models.

821
00:29:04,000 --> 00:29:05,640
It is AI self-sufficiency.

822
00:29:05,640 --> 00:29:07,160
Microsoft wants to reach a point

823
00:29:07,160 --> 00:29:09,800
where they are not permanently renting someone else's compute

824
00:29:09,800 --> 00:29:11,640
to run their own intelligence layer.

825
00:29:11,640 --> 00:29:14,200
Owning the chip, the model and the relationship between them

826
00:29:14,200 --> 00:29:17,240
is what makes self-sufficiency more than a talking point.

827
00:29:17,240 --> 00:29:19,840
On paper, this is a genuinely elegant design.

828
00:29:19,840 --> 00:29:22,040
You have sparse activation for cost control,

829
00:29:22,040 --> 00:29:23,760
custom silicon for efficiency

830
00:29:23,760 --> 00:29:26,000
and a strategic goal tying it all together.

831
00:29:26,000 --> 00:29:28,800
But an architecture diagram is not proof of anything.

832
00:29:28,800 --> 00:29:29,920
None of this actually matters

833
00:29:29,920 --> 00:29:32,280
unless it holds up outside a keynote slide.

834
00:29:32,280 --> 00:29:34,280
Under real workloads with real deadlines

835
00:29:34,280 --> 00:29:37,680
and real budgets on the line, the Excel proof point.

836
00:29:37,680 --> 00:29:40,200
So here is what all of that architecture actually produced

837
00:29:40,200 --> 00:29:42,360
when Microsoft pointed it at a real problem

838
00:29:42,360 --> 00:29:44,120
instead of a benchmark chart.

839
00:29:44,120 --> 00:29:46,160
Microsoft wanted agentic Excel tasks

840
00:29:46,160 --> 00:29:47,960
that actually held up under real use.

841
00:29:47,960 --> 00:29:50,440
They didn't want a demo where the model reads a spreadsheet

842
00:29:50,440 --> 00:29:53,720
and summarizes it once for an audience that claps and moves on.

843
00:29:53,720 --> 00:29:55,560
They needed real agentic work.

844
00:29:55,560 --> 00:29:57,600
The kind where a model has to look at a workbook

845
00:29:57,600 --> 00:29:59,600
understand what is actually being asked,

846
00:29:59,600 --> 00:30:02,400
take multiple steps and get it right every single time.

847
00:30:02,400 --> 00:30:04,400
And it has to do that across the massive volume

848
00:30:04,400 --> 00:30:07,720
of requests an actual Excel user-based generates every day.

849
00:30:07,720 --> 00:30:08,800
But here is the problem.

850
00:30:08,800 --> 00:30:11,440
General purpose frontier models could technically do this.

851
00:30:11,440 --> 00:30:14,120
But at the volume Excel operates at hundreds of millions

852
00:30:14,120 --> 00:30:17,000
of users and an enormous number of daily requests.

853
00:30:17,000 --> 00:30:19,240
Those models were either too slow to feel usable

854
00:30:19,240 --> 00:30:21,000
or too expensive to run at scale

855
00:30:21,000 --> 00:30:22,640
without the economics falling apart.

856
00:30:22,640 --> 00:30:25,000
You can build an impressive demo with a frontier model

857
00:30:25,000 --> 00:30:26,840
answering one spreadsheet question.

858
00:30:26,840 --> 00:30:28,720
But you cannot afford to run that same model

859
00:30:28,720 --> 00:30:30,920
at that same depth for every formula question

860
00:30:30,920 --> 00:30:33,080
and data cleanup task hitting Excel in production.

861
00:30:33,080 --> 00:30:35,360
So this is the exact gap we have been building toward.

862
00:30:35,360 --> 00:30:36,680
This isn't a hypothetical.

863
00:30:36,680 --> 00:30:39,440
It is an actual product decision Microsoft had to make.

864
00:30:39,440 --> 00:30:40,520
And here is the shift.

865
00:30:40,520 --> 00:30:44,600
M.A.I. tuned models ended up matching GPT 5.4 on the relevant bench

866
00:30:44,600 --> 00:30:47,200
marks while running a 10 times greater cost efficiency.

867
00:30:47,200 --> 00:30:48,200
Read that again.

868
00:30:48,200 --> 00:30:50,200
Because it is the whole thesis of this episode

869
00:30:50,200 --> 00:30:53,200
compressed into one result, it wasn't close enough performance

870
00:30:53,200 --> 00:30:54,440
for a discount.

871
00:30:54,440 --> 00:30:56,680
It was matching performance at a 10th of the cost.

872
00:30:56,680 --> 00:30:57,760
That is not a trade-off.

873
00:30:57,760 --> 00:31:00,200
That is what happens when a model gets tuned specifically

874
00:31:00,200 --> 00:31:03,120
for a task instead of being asked to be brilliant at everything.

875
00:31:03,120 --> 00:31:04,360
And this was not a one-off.

876
00:31:04,360 --> 00:31:07,480
The same pattern repeated with McKinsey's task-specific tuning.

877
00:31:07,480 --> 00:31:10,160
When M.A.I. got tuned on McKinsey's actual workflows,

878
00:31:10,160 --> 00:31:13,000
it delivered the highest win rate of any model tested and out

879
00:31:13,000 --> 00:31:16,320
performed GPT 5.5 on their specific tasks.

880
00:31:16,320 --> 00:31:18,560
It landed at that same 10 times efficiency gain

881
00:31:18,560 --> 00:31:20,280
to completely different organizations,

882
00:31:20,280 --> 00:31:21,720
two completely different workloads.

883
00:31:21,720 --> 00:31:23,800
Spreadsheets on one side and consulting workflows

884
00:31:23,800 --> 00:31:24,480
on the other.

885
00:31:24,480 --> 00:31:26,080
The same result shows up both times.

886
00:31:26,080 --> 00:31:28,240
So what is actually happening is a pattern.

887
00:31:28,240 --> 00:31:30,320
One win could be a fluke or a bench mark

888
00:31:30,320 --> 00:31:32,480
that happened to favor M.A.I.I's training data.

889
00:31:32,480 --> 00:31:34,240
But two wins in unrelated domains

890
00:31:34,240 --> 00:31:36,520
landing on the same 10 times efficiency number

891
00:31:36,520 --> 00:31:37,800
means something.

892
00:31:37,800 --> 00:31:40,960
It is a sign the gain isn't specific to spreadsheets or consulting.

893
00:31:40,960 --> 00:31:42,920
It is what shows up whenever reasoning and runtime

894
00:31:42,920 --> 00:31:45,200
get tuned together against a real narrow workflow

895
00:31:45,200 --> 00:31:47,280
instead of being asked to perform generally.

896
00:31:47,280 --> 00:31:48,600
And that is the real point here.

897
00:31:48,600 --> 00:31:50,120
This is not a demo.

898
00:31:50,120 --> 00:31:52,520
Demo's show you what is possible in ideal conditions.

899
00:31:52,520 --> 00:31:55,080
This is what happens when the architecture we have described

900
00:31:55,080 --> 00:31:57,120
gets pointed at actual production workloads

901
00:31:57,120 --> 00:31:59,520
with actual volume and actual cost pressure.

902
00:31:59,520 --> 00:32:02,000
These are real users who would notice if it broke,

903
00:32:02,000 --> 00:32:03,680
which raises a question worth sitting with.

904
00:32:03,680 --> 00:32:05,400
If tuning a model on your own workflows

905
00:32:05,400 --> 00:32:06,920
gets your results like this.

906
00:32:06,920 --> 00:32:08,800
What does that actually mean for who ends up owning

907
00:32:08,800 --> 00:32:11,240
the model that comes out the other side?

908
00:32:11,240 --> 00:32:13,880
Frontier tuning and wired changes who owns the model.

909
00:32:13,880 --> 00:32:15,840
The answer starts with a piece of infrastructure

910
00:32:15,840 --> 00:32:18,320
Microsoft calls reinforcement learning environments.

911
00:32:18,320 --> 00:32:19,520
Or RLEs.

912
00:32:19,520 --> 00:32:21,600
Think of an RLE as a training gym built

913
00:32:21,600 --> 00:32:23,200
for exactly one company's problems.

914
00:32:23,200 --> 00:32:25,040
It is not a general gym with treadmills

915
00:32:25,040 --> 00:32:26,520
for anybody off the street.

916
00:32:26,520 --> 00:32:29,360
It is a facility custom built around the specific movements

917
00:32:29,360 --> 00:32:31,600
one team actually needs to get good at.

918
00:32:31,600 --> 00:32:34,520
Microsoft used its own RLEs combined with M.A.I.I models

919
00:32:34,520 --> 00:32:36,680
to climb to what better performance on Excel tasks

920
00:32:36,680 --> 00:32:37,680
specifically.

921
00:32:37,680 --> 00:32:39,720
McKinsey did the same thing with their own tasks

922
00:32:39,720 --> 00:32:41,040
in their own environment.

923
00:32:41,040 --> 00:32:42,240
The gym isn't shared.

924
00:32:42,240 --> 00:32:44,720
It is built around one organization's actual work.

925
00:32:44,720 --> 00:32:46,480
And the model only gets strong at the things

926
00:32:46,480 --> 00:32:47,680
that gym trains it on.

927
00:32:47,680 --> 00:32:49,440
But here is the differentiator that matters more

928
00:32:49,440 --> 00:32:50,480
than anything else.

929
00:32:50,480 --> 00:32:52,240
You don't rent shared intelligence.

930
00:32:52,240 --> 00:32:53,400
You keep the resulting model.

931
00:32:53,400 --> 00:32:54,520
Sit with that for a second.

932
00:32:54,520 --> 00:32:56,400
Because it is a genuinely different relationship

933
00:32:56,400 --> 00:32:58,480
than most companies have with A.I.I right now.

934
00:32:58,480 --> 00:33:00,640
Most SAS A.I.Tools work the opposite way.

935
00:33:00,640 --> 00:33:01,440
You use the tool.

936
00:33:01,440 --> 00:33:03,320
Your usage feeds a shared system.

937
00:33:03,320 --> 00:33:05,400
And every improvement that comes from your data

938
00:33:05,400 --> 00:33:07,160
gets folded into the same general model

939
00:33:07,160 --> 00:33:09,240
every other customer is also using.

940
00:33:09,240 --> 00:33:11,480
Your competitor uses the same tool and benefits

941
00:33:11,480 --> 00:33:14,240
from the exact same improvements your usage helped create.

942
00:33:14,240 --> 00:33:15,960
You are not building anything proprietary.

943
00:33:15,960 --> 00:33:17,880
You are contributing to somebody else's product

944
00:33:17,880 --> 00:33:19,240
one query at a time.

945
00:33:19,240 --> 00:33:21,840
And you don't get to walk away with what you helped build.

946
00:33:21,840 --> 00:33:23,360
Frontier tuning breaks that arrangement.

947
00:33:23,360 --> 00:33:25,840
When you tune M.A.I.I models inside your own RLE

948
00:33:25,840 --> 00:33:28,280
on your own tasks, the resulting model is yours.

949
00:33:28,280 --> 00:33:30,600
It is not a shared checkpoint everyone draws from.

950
00:33:30,600 --> 00:33:33,000
It is a model shaped specifically by your workflows

951
00:33:33,000 --> 00:33:34,280
and your institutional know-how.

952
00:33:34,280 --> 00:33:36,800
Nobody else gets to touch what came out of that process.

953
00:33:36,800 --> 00:33:39,000
And this is the shift where your workflows become your mode

954
00:33:39,000 --> 00:33:40,240
instead of a vendor's mode.

955
00:33:40,240 --> 00:33:42,800
Every company has some version of tribal knowledge.

956
00:33:42,800 --> 00:33:45,120
The specific way they handle a certain kind of customer

957
00:33:45,120 --> 00:33:47,000
complaint or the particular judgment calls

958
00:33:47,000 --> 00:33:49,600
that separate a senior employee from a junior one.

959
00:33:49,600 --> 00:33:52,240
Normally that knowledge stays locked inside people's heads

960
00:33:52,240 --> 00:33:54,280
or scattered across documentation.

961
00:33:54,280 --> 00:33:55,200
Nobody reads.

962
00:33:55,200 --> 00:33:56,840
Frontier tuning turns that same knowledge

963
00:33:56,840 --> 00:33:58,280
into a model asset.

964
00:33:58,280 --> 00:34:00,560
The workflows that make your company good at what it does

965
00:34:00,560 --> 00:34:02,360
become the thing the model gets tuned on.

966
00:34:02,360 --> 00:34:04,240
The resulting intelligence belongs to you.

967
00:34:04,240 --> 00:34:06,240
Not to whoever built the underlying model.

968
00:34:06,240 --> 00:34:08,200
Now connect this back to the architecture we have been

969
00:34:08,200 --> 00:34:09,000
describing.

970
00:34:09,000 --> 00:34:10,840
Phi 4 becomes the tunable runtime.

971
00:34:10,840 --> 00:34:12,880
The fast and local layer that can be shaped around your

972
00:34:12,880 --> 00:34:14,520
specific execution needs.

973
00:34:14,520 --> 00:34:17,360
M.I.I.I.I sits underneath as the reasoning foundation.

974
00:34:17,360 --> 00:34:20,760
The deep layer custom agents draw on when a request actually

975
00:34:20,760 --> 00:34:22,800
needs planning instead of just action.

976
00:34:22,800 --> 00:34:25,320
Frontier tuning is what lets both layers get shaped around

977
00:34:25,320 --> 00:34:28,600
one company's actual work instead of staying generic and shared

978
00:34:28,600 --> 00:34:31,760
across every customer using the same off-the-shelf model.

979
00:34:31,760 --> 00:34:33,800
And this is where things change for IT decision-makers

980
00:34:33,800 --> 00:34:36,920
up until now choosing an AI vendor mostly meant choosing a tool.

981
00:34:36,920 --> 00:34:39,280
Now it means choosing whether the intelligence your company

982
00:34:39,280 --> 00:34:41,520
builds through daily use stays yours

983
00:34:41,520 --> 00:34:44,360
or quietly becomes part of somebody else's shared product.

984
00:34:44,360 --> 00:34:46,680
That is not a question engineering teams have historically

985
00:34:46,680 --> 00:34:49,000
had to ask, but it is about to become one of the most

986
00:34:49,000 --> 00:34:50,720
consequential calls they make.

987
00:34:50,720 --> 00:34:53,520
If this changed how you think, follow me, Mirko Peters,

988
00:34:53,520 --> 00:34:54,680
on LinkedIn.

989
00:34:54,680 --> 00:34:56,680
And if you want more of this, leave a review.

990
00:34:56,680 --> 00:34:57,880
It helps more people find it.

991
00:34:57,880 --> 00:35:00,400
Share this with your team, especially if you are dealing

992
00:35:00,400 --> 00:35:01,440
with this right now.

993
00:35:01,440 --> 00:35:03,800
What this means for IT architects and admins.

994
00:35:03,800 --> 00:35:05,680
So here is where this lands for the people who actually

995
00:35:05,680 --> 00:35:06,680
have to build it.

996
00:35:06,680 --> 00:35:09,360
For years, the standard IT question was simple.

997
00:35:09,360 --> 00:35:11,000
Which model should we standardize on?

998
00:35:11,000 --> 00:35:12,680
You pick the vendor, you pick the model,

999
00:35:12,680 --> 00:35:15,200
you roll it out across the org, and then everyone

1000
00:35:15,200 --> 00:35:16,480
built against that one choice.

1001
00:35:16,480 --> 00:35:18,880
That question made sense when there was only one kind of model

1002
00:35:18,880 --> 00:35:19,760
to pick from.

1003
00:35:19,760 --> 00:35:21,160
It does not make sense anymore.

1004
00:35:21,160 --> 00:35:23,480
Standardizing on a single model in an architecture

1005
00:35:23,480 --> 00:35:25,560
built around a reason and runtime split

1006
00:35:25,560 --> 00:35:28,600
is like hiring one person to do every job in your company.

1007
00:35:28,600 --> 00:35:30,680
You wouldn't ask the same person to file paperwork

1008
00:35:30,680 --> 00:35:32,000
and negotiate a merger.

1009
00:35:32,000 --> 00:35:34,440
The question was never wrong because the answer was hard.

1010
00:35:34,440 --> 00:35:36,960
It was wrong because it assumed the wrong thing needed

1011
00:35:36,960 --> 00:35:37,760
choosing.

1012
00:35:37,760 --> 00:35:39,160
The actual question is different.

1013
00:35:39,160 --> 00:35:40,400
What is our rooting logic?

1014
00:35:40,400 --> 00:35:41,720
And who owns the decision layer?

1015
00:35:41,720 --> 00:35:43,560
That is the only question that matters now.

1016
00:35:43,560 --> 00:35:46,040
Not which model, but who decides which model?

1017
00:35:46,040 --> 00:35:47,200
And based on what?

1018
00:35:47,200 --> 00:35:48,960
Because as we saw, the routing layer

1019
00:35:48,960 --> 00:35:51,560
is the part nobody wants to build and everybody needs.

1020
00:35:51,560 --> 00:35:54,000
Someone in the organization has to own that decision layer.

1021
00:35:54,000 --> 00:35:56,320
Someone has to be responsible for the classification logic

1022
00:35:56,320 --> 00:35:58,520
that sends a routine request to 5/4

1023
00:35:58,520 --> 00:36:02,320
and a genuinely hard one up to my one right now in most organizations.

1024
00:36:02,320 --> 00:36:03,320
Nobody owns that.

1025
00:36:03,320 --> 00:36:06,080
It is nobody's job, which means it is not getting built,

1026
00:36:06,080 --> 00:36:09,640
which means every request still defaults to whatever single model

1027
00:36:09,640 --> 00:36:11,080
happens to be plugged in.

1028
00:36:11,080 --> 00:36:14,520
That ownership gap points to a skill most IT teams do not have yet.

1029
00:36:14,520 --> 00:36:18,080
Workload classification, not model tuning, not prompt engineering,

1030
00:36:18,080 --> 00:36:19,560
workload classification.

1031
00:36:19,560 --> 00:36:21,880
This is the ability to look at an incoming request

1032
00:36:21,880 --> 00:36:24,560
and judge whether it is simple or reasoning heavy.

1033
00:36:24,560 --> 00:36:27,720
Before it ever reaches a model, that is a genuinely new skill.

1034
00:36:27,720 --> 00:36:29,480
Most architects were never trained to do this,

1035
00:36:29,480 --> 00:36:31,680
because until recently, there was no reason to do it.

1036
00:36:31,680 --> 00:36:34,040
Every request went to the same place regardless.

1037
00:36:34,040 --> 00:36:35,960
Now, that distinction is the whole game.

1038
00:36:35,960 --> 00:36:37,400
And most teams are starting from zero.

1039
00:36:37,400 --> 00:36:40,160
There is a governance layer sitting on top of all this too.

1040
00:36:40,160 --> 00:36:42,200
On-prem 5/4 deployments become the answer

1041
00:36:42,200 --> 00:36:44,240
when data sovereignty is non-negotiable,

1042
00:36:44,240 --> 00:36:46,240
when a request can never leave the building,

1043
00:36:46,240 --> 00:36:47,720
or the endpoint, or the country.

1044
00:36:47,720 --> 00:36:48,920
You go local.

1045
00:36:48,920 --> 00:36:50,800
Cloud-based MI1 becomes the answer

1046
00:36:50,800 --> 00:36:52,200
when you need centralized reasoning,

1047
00:36:52,200 --> 00:36:54,760
when you need shared context across a whole organization,

1048
00:36:54,760 --> 00:36:56,840
or the kind of planning work that benefits from sitting

1049
00:36:56,840 --> 00:36:57,680
in one place.

1050
00:36:57,680 --> 00:36:58,680
Go to the cloud.

1051
00:36:58,680 --> 00:37:00,520
That is not a technology decision anymore.

1052
00:37:00,520 --> 00:37:02,280
That is a governance decision.

1053
00:37:02,280 --> 00:37:04,680
And it has to be made deliberately, not by default.

1054
00:37:04,680 --> 00:37:07,440
And here is the real friction point, stated plainly.

1055
00:37:07,440 --> 00:37:10,400
Most organizations do not have this rooting infrastructure yet.

1056
00:37:10,400 --> 00:37:11,240
It does not exist.

1057
00:37:11,240 --> 00:37:12,040
The models exist.

1058
00:37:12,040 --> 00:37:13,040
5/4 is sitting there.

1059
00:37:13,040 --> 00:37:13,960
MIT licensed.

1060
00:37:13,960 --> 00:37:14,880
Ready to deploy.

1061
00:37:14,880 --> 00:37:16,720
MI1 is available through Foundry.

1062
00:37:16,720 --> 00:37:19,280
But the layer that decides which request goes where?

1063
00:37:19,280 --> 00:37:22,560
The classification logic, the ownership, the governance rules.

1064
00:37:22,560 --> 00:37:24,160
That is the actual work still ahead.

1065
00:37:24,160 --> 00:37:25,520
Nobody sells you that off the shelf.

1066
00:37:25,520 --> 00:37:27,280
You have to build it, which is a very different

1067
00:37:27,280 --> 00:37:29,440
conversation depending on who is sitting in the room

1068
00:37:29,440 --> 00:37:30,760
for architects and admins.

1069
00:37:30,760 --> 00:37:32,320
This is an infrastructure problem.

1070
00:37:32,320 --> 00:37:34,000
For the people signing the budget,

1071
00:37:34,000 --> 00:37:36,480
it looks like something else entirely.

1072
00:37:36,480 --> 00:37:39,000
What this means for decision makers and budget owners.

1073
00:37:39,000 --> 00:37:40,560
Here is how that conversation changes

1074
00:37:40,560 --> 00:37:42,960
once it reaches whoever signs the check.

1075
00:37:42,960 --> 00:37:46,400
Most executives still frame this as, how much does AI cost?

1076
00:37:46,400 --> 00:37:47,280
The wrong question.

1077
00:37:47,280 --> 00:37:49,760
The real question is, how much mis-rooted AI costs?

1078
00:37:49,760 --> 00:37:51,720
Because that is the number nobody is tracking.

1079
00:37:51,720 --> 00:37:53,640
It is not the invoice from the model provider.

1080
00:37:53,640 --> 00:37:56,200
It is the invisible cost sitting underneath it.

1081
00:37:56,200 --> 00:37:59,000
Every request that got sent to an expensive model

1082
00:37:59,000 --> 00:38:01,160
when a cheap one would have handled it just as well.

1083
00:38:01,160 --> 00:38:01,960
Is a loss.

1084
00:38:01,960 --> 00:38:04,000
That is not a line item on any budget report.

1085
00:38:04,000 --> 00:38:05,040
It is a leak.

1086
00:38:05,040 --> 00:38:07,280
And leaks do not show up until someone finally goes looking

1087
00:38:07,280 --> 00:38:07,760
for them.

1088
00:38:07,760 --> 00:38:09,040
Look at what correct rooting actually

1089
00:38:09,040 --> 00:38:10,400
produces in practice.

1090
00:38:10,400 --> 00:38:12,000
In one case, fuel cost reduction

1091
00:38:12,000 --> 00:38:14,320
was tied directly to model routing decisions.

1092
00:38:14,320 --> 00:38:16,840
Those savings showed up because the right model was matched

1093
00:38:16,840 --> 00:38:18,880
to the right task at the right layer.

1094
00:38:18,880 --> 00:38:20,880
We see the same pattern on the support side.

1095
00:38:20,880 --> 00:38:23,800
Latency was cut from 15 seconds down to under two.

1096
00:38:23,800 --> 00:38:26,440
That drop did not come from a faster frontier model.

1097
00:38:26,440 --> 00:38:28,440
It came from rooting routine requests,

1098
00:38:28,440 --> 00:38:31,280
somewhere that never needed frontier depth in the first place.

1099
00:38:31,280 --> 00:38:33,360
That is the whole argument made concrete.

1100
00:38:33,360 --> 00:38:35,000
The game was not capability.

1101
00:38:35,000 --> 00:38:36,400
It was placement.

1102
00:38:36,400 --> 00:38:38,640
Now, here is the number that turns this

1103
00:38:38,640 --> 00:38:40,400
into a real financial decision instead

1104
00:38:40,400 --> 00:38:41,680
of a technical preference.

1105
00:38:41,680 --> 00:38:43,480
Self-hosted small models break even

1106
00:38:43,480 --> 00:38:46,800
against managed frontier APIs inside 18 months.

1107
00:38:46,800 --> 00:38:48,920
Once you are operating at volume, 18 months

1108
00:38:48,920 --> 00:38:51,320
that is not a speculative payback period stretched out

1109
00:38:51,320 --> 00:38:53,880
over a decade to make the math look better on a slide.

1110
00:38:53,880 --> 00:38:56,520
That is a timeline a CFO can actually plan around.

1111
00:38:56,520 --> 00:38:58,640
At real volume, the math does not stay close.

1112
00:38:58,640 --> 00:39:00,600
It tilts hard toward whichever company

1113
00:39:00,600 --> 00:39:02,440
built the rooting infrastructure early,

1114
00:39:02,440 --> 00:39:04,440
which is exactly why this stops being purely

1115
00:39:04,440 --> 00:39:07,200
an engineering decision and becomes a board level conversation.

1116
00:39:07,200 --> 00:39:10,200
Engineering teams can debate architecture patterns all day,

1117
00:39:10,200 --> 00:39:13,240
but an 18 month break even on infrastructure spend.

1118
00:39:13,240 --> 00:39:17,000
Tied to a measurable latency and cost outcome is a different story.

1119
00:39:17,000 --> 00:39:19,760
That is the kind of number that gets a line in a quarterly review.

1120
00:39:19,760 --> 00:39:22,680
Once the payback period is that short and that provable,

1121
00:39:22,680 --> 00:39:24,520
the decision is not, should we let the engineers

1122
00:39:24,520 --> 00:39:25,560
experiment with this?

1123
00:39:25,560 --> 00:39:28,840
A, the decision is, why haven't we already funded it?

1124
00:39:28,840 --> 00:39:30,840
And here are the stakes worth sitting with.

1125
00:39:30,840 --> 00:39:33,800
Because it does not stay contained to one budget cycle.

1126
00:39:33,800 --> 00:39:35,320
The companies that get routing right

1127
00:39:35,320 --> 00:39:37,920
are not just saving money on this quarter's cloud bill.

1128
00:39:37,920 --> 00:39:40,320
They are operating at a structurally lower cost basis

1129
00:39:40,320 --> 00:39:42,080
than competitors who never built the split.

1130
00:39:42,080 --> 00:39:43,520
That is not a temporary advantage

1131
00:39:43,520 --> 00:39:45,400
that erodes once everyone catches up.

1132
00:39:45,400 --> 00:39:46,320
It compounds.

1133
00:39:46,320 --> 00:39:48,680
Every request processed at the correct layer.

1134
00:39:48,680 --> 00:39:51,440
Instead of defaulting to the most expensive option available,

1135
00:39:51,440 --> 00:39:53,200
is margin the other company does not have.

1136
00:39:53,200 --> 00:39:55,120
It stays that way, quarter after quarter,

1137
00:39:55,120 --> 00:39:56,960
at whatever scale they are operating at.

1138
00:39:56,960 --> 00:39:59,040
So the case for budget owners is straightforward.

1139
00:39:59,040 --> 00:40:01,040
This is not a bet on which model wins.

1140
00:40:01,040 --> 00:40:02,880
It is a bet on whether your cost structure

1141
00:40:02,880 --> 00:40:04,400
looks like the company is still routing

1142
00:40:04,400 --> 00:40:06,240
everything through one expensive door,

1143
00:40:06,240 --> 00:40:08,920
or the ones who build the infrastructure to stop doing that.

1144
00:40:08,920 --> 00:40:11,200
But before this turns into an uncritical pitch,

1145
00:40:11,200 --> 00:40:13,200
there is a tension worth being honest about,

1146
00:40:13,200 --> 00:40:16,520
one that sits uncomfortably next to everything just described.

1147
00:40:16,520 --> 00:40:18,800
The trust gap Microsoft hasn't solved yet.

1148
00:40:18,800 --> 00:40:20,000
There is a tension here.

1149
00:40:20,000 --> 00:40:23,280
It is worth saying out loud, instead of skipping past it.

1150
00:40:23,280 --> 00:40:25,920
Microsoft pitches Copilot as a serious enterprise tool.

1151
00:40:25,920 --> 00:40:26,960
They wire it into Excel.

1152
00:40:26,960 --> 00:40:27,920
They put it in Teams.

1153
00:40:27,920 --> 00:40:29,680
They tell you it belongs in the exact workflows

1154
00:40:29,680 --> 00:40:30,880
we have been talking about.

1155
00:40:30,880 --> 00:40:32,480
But then you look at the fine print.

1156
00:40:32,480 --> 00:40:35,520
Copilot's own terms of use still describe the product

1157
00:40:35,520 --> 00:40:38,040
as being for entertainment purposes only.

1158
00:40:38,040 --> 00:40:39,400
That is the actual language.

1159
00:40:39,400 --> 00:40:40,760
It is not a marketing footnote.

1160
00:40:40,760 --> 00:40:42,360
It is a legal disclaimer telling you

1161
00:40:42,360 --> 00:40:44,400
not to treat the outputs as reliable.

1162
00:40:44,400 --> 00:40:45,720
And Microsoft is not alone.

1163
00:40:45,720 --> 00:40:48,480
Open AI and XAI have versions of the same language

1164
00:40:48,480 --> 00:40:49,400
in their own terms.

1165
00:40:49,400 --> 00:40:50,920
This is an industry-wide pattern.

1166
00:40:50,920 --> 00:40:52,520
Corsion is baked into the legal layer

1167
00:40:52,520 --> 00:40:54,360
of every major AI provider.

1168
00:40:54,360 --> 00:40:56,280
It does not matter how confidently they talk

1169
00:40:56,280 --> 00:40:58,040
about the technology in public.

1170
00:40:58,040 --> 00:40:59,320
So here's the problem.

1171
00:40:59,320 --> 00:41:01,880
These companies want AI to be critical infrastructure.

1172
00:41:01,880 --> 00:41:04,160
They wanted to be the engine behind your Excel agents

1173
00:41:04,160 --> 00:41:05,160
and your support systems.

1174
00:41:05,160 --> 00:41:06,400
They wanted to handle decisions

1175
00:41:06,400 --> 00:41:08,520
that affect real budgets and real customers.

1176
00:41:08,520 --> 00:41:10,760
But at the same time, the legal language says,

1177
00:41:10,760 --> 00:41:12,360
don't actually rely on this.

1178
00:41:12,360 --> 00:41:13,840
You have critical infrastructure

1179
00:41:13,840 --> 00:41:16,400
and legally-disclamed unreliability sitting

1180
00:41:16,400 --> 00:41:18,080
in the same product at the same time.

1181
00:41:18,080 --> 00:41:19,560
That gap matters if you are planning

1182
00:41:19,560 --> 00:41:20,880
to actually use this in a business.

1183
00:41:20,880 --> 00:41:23,120
It is not a reason to avoid the architecture we have been

1184
00:41:23,120 --> 00:41:25,880
building, but it is a reason not to build dependencies

1185
00:41:25,880 --> 00:41:27,280
that you cannot reverse.

1186
00:41:27,280 --> 00:41:30,000
Nobody is willing to formally stand behind these outputs yet.

1187
00:41:30,000 --> 00:41:32,800
If a workflow can handle an occasional wrong answer,

1188
00:41:32,800 --> 00:41:33,760
that is one thing.

1189
00:41:33,760 --> 00:41:34,440
You catch it.

1190
00:41:34,440 --> 00:41:35,280
You correct it.

1191
00:41:35,280 --> 00:41:37,720
You route it to a human when confidence is low.

1192
00:41:37,720 --> 00:41:40,640
But if your workflow assumes the model is simply correct,

1193
00:41:40,640 --> 00:41:41,920
you are taking a massive risk.

1194
00:41:41,920 --> 00:41:43,240
That is the risk the terms of use

1195
00:41:43,240 --> 00:41:45,400
are warning you about whether you read them or not.

1196
00:41:45,400 --> 00:41:47,440
Now, we should add some context here.

1197
00:41:47,440 --> 00:41:50,680
Microsoft has acknowledged that this wording is outdated.

1198
00:41:50,680 --> 00:41:51,600
That matters.

1199
00:41:51,600 --> 00:41:53,600
A company saying the language does not reflect

1200
00:41:53,600 --> 00:41:56,200
the product is different from a company defending it

1201
00:41:56,200 --> 00:41:57,120
as accurate.

1202
00:41:57,120 --> 00:41:59,000
It suggests this is just a lagging artifact.

1203
00:41:59,000 --> 00:42:01,960
It was written years before Copilot could do what it does today.

1204
00:42:01,960 --> 00:42:03,520
But acknowledging outdated language

1205
00:42:03,520 --> 00:42:05,720
and actually fixing the gap are two different things.

1206
00:42:05,720 --> 00:42:07,200
Only one of those has happened.

1207
00:42:07,200 --> 00:42:09,280
This does not change the strategy we have covered.

1208
00:42:09,280 --> 00:42:10,960
The reason and runtime split still works.

1209
00:42:10,960 --> 00:42:13,400
The routing logic and the cost math still hold up.

1210
00:42:13,400 --> 00:42:15,040
What it means is that governance has

1211
00:42:15,040 --> 00:42:16,400
to catch up to the technology.

1212
00:42:16,400 --> 00:42:19,160
The architecture is ahead of the paperwork right now.

1213
00:42:19,160 --> 00:42:20,840
And until that paperwork catches up,

1214
00:42:20,840 --> 00:42:23,160
the responsible move is to build verification

1215
00:42:23,160 --> 00:42:24,040
into your workflow.

1216
00:42:24,040 --> 00:42:26,120
Do not assume the legal language will update itself

1217
00:42:26,120 --> 00:42:28,840
before something important breaks.

1218
00:42:28,840 --> 00:42:31,360
Multimodality as the next layer of the runtime.

1219
00:42:31,360 --> 00:42:33,400
Now zoom out from governance for a second.

1220
00:42:33,400 --> 00:42:35,760
There is another layer to this runtime story.

1221
00:42:35,760 --> 00:42:37,240
It lives inside the word multimodal,

1222
00:42:37,240 --> 00:42:39,280
whether when you look at 5 or 4 multimodal,

1223
00:42:39,280 --> 00:42:42,640
it does not run vision, audio and text as three separate products.

1224
00:42:42,640 --> 00:42:45,240
It is not a text model with a camera and a microphone

1225
00:42:45,240 --> 00:42:46,760
bolted on as afterthoughts.

1226
00:42:46,760 --> 00:42:49,640
Those setups have separate pipelines and separate failure points.

1227
00:42:49,640 --> 00:42:50,760
This is one backbone.

1228
00:42:50,760 --> 00:42:53,360
One core handles all three kinds of input.

1229
00:42:53,360 --> 00:42:55,360
Here is what that looks like in the real world.

1230
00:42:55,360 --> 00:42:59,080
You get OCR accuracy that holds up on messy documents.

1231
00:42:59,080 --> 00:43:00,880
You can hand the model an image and ask

1232
00:43:00,880 --> 00:43:02,920
specific questions about what is in it.

1233
00:43:02,920 --> 00:43:05,120
You get answers grounded in what is actually there.

1234
00:43:05,120 --> 00:43:08,160
You get audio transcription and translation in one single pass.

1235
00:43:08,160 --> 00:43:11,200
There is no separate translation step added on afterward.

1236
00:43:11,200 --> 00:43:14,640
It is one model, one forward pass, three different senses

1237
00:43:14,640 --> 00:43:16,280
feeding into the same understanding.

1238
00:43:16,280 --> 00:43:17,800
Why does this matter for developers?

1239
00:43:17,800 --> 00:43:19,640
Because the alternative is a nightmare.

1240
00:43:19,640 --> 00:43:22,160
Separate pipelines mean separate failure points

1241
00:43:22,160 --> 00:43:23,800
and separate latency budgets.

1242
00:43:23,800 --> 00:43:26,040
You have to keep different versions synchronized.

1243
00:43:26,040 --> 00:43:28,040
Things break quietly and you do not notice

1244
00:43:28,040 --> 00:43:29,280
until a user complains.

1245
00:43:29,280 --> 00:43:31,160
A unified backbone collapses all of that.

1246
00:43:31,160 --> 00:43:33,480
It gives you one thing to deploy, one thing to monitor,

1247
00:43:33,480 --> 00:43:34,680
one thing to update.

1248
00:43:34,680 --> 00:43:36,720
That is the difference between a runtime you can actually

1249
00:43:36,720 --> 00:43:40,800
maintain and one that becomes unmanageable as you add more inputs.

1250
00:43:40,800 --> 00:43:42,600
This pattern shows up on the MAI side too.

1251
00:43:42,600 --> 00:43:45,920
They use different names like MAI voice or MI image.

1252
00:43:45,920 --> 00:43:48,240
But they follow the exact same logic we have been describing.

1253
00:43:48,240 --> 00:43:49,840
Reasoning heavy work sits in one place.

1254
00:43:49,840 --> 00:43:52,240
Fast specialized execution sits somewhere else.

1255
00:43:52,240 --> 00:43:54,440
It stays closer to where the input is happening.

1256
00:43:54,440 --> 00:43:57,040
It is easy to think multi-modal is a separate category.

1257
00:43:57,040 --> 00:43:57,680
It isn't.

1258
00:43:57,680 --> 00:44:00,120
Even inside these systems, the same split holds.

1259
00:44:00,120 --> 00:44:02,840
Heavy reasoning needs deep context and multi-step planning

1260
00:44:02,840 --> 00:44:04,480
that stays with the reasoning layer.

1261
00:44:04,480 --> 00:44:07,160
Fast execution needs to transcribe audio or read text

1262
00:44:07,160 --> 00:44:08,560
off a scan in real time.

1263
00:44:08,560 --> 00:44:10,200
That stays with the runtime layer.

1264
00:44:10,200 --> 00:44:12,720
Adding vision and audio did not break the reason and runtime

1265
00:44:12,720 --> 00:44:13,400
pattern.

1266
00:44:13,400 --> 00:44:16,040
It just gave the pattern more kinds of input to sort through.

1267
00:44:16,040 --> 00:44:18,200
The architecture does not get more complicated

1268
00:44:18,200 --> 00:44:19,720
as you add modalities.

1269
00:44:19,720 --> 00:44:23,000
It just gets applied more broadly, which leads to the next question.

1270
00:44:23,000 --> 00:44:25,680
If this split holds true across text, vision, and audio

1271
00:44:25,680 --> 00:44:27,720
today, where does Microsoft take it next?

1272
00:44:27,720 --> 00:44:30,360
And what does that tell you about what is coming?

1273
00:44:30,360 --> 00:44:32,800
The trajectory where this architecture goes next.

1274
00:44:32,800 --> 00:44:35,240
Microsoft is signaling where this is headed.

1275
00:44:35,240 --> 00:44:38,040
And it is much bigger than the seven models we saw at build.

1276
00:44:38,040 --> 00:44:40,520
The plan does not stop at task-specific models.

1277
00:44:40,520 --> 00:44:43,320
Mustafa Saliman has been very direct about the goal here.

1278
00:44:43,320 --> 00:44:45,520
They are building full-frontier LLMs

1279
00:44:45,520 --> 00:44:47,640
to compete with systems like GPT-4.

1280
00:44:47,640 --> 00:44:49,320
They aren't just building narrow tools

1281
00:44:49,320 --> 00:44:50,960
for coding or transcription.

1282
00:44:50,960 --> 00:44:52,280
Everything we have looked at so far,

1283
00:44:52,280 --> 00:44:54,800
the MAI thinking models, the Excel tuning,

1284
00:44:54,800 --> 00:44:57,680
the McKinsey results, that is just the current chapter.

1285
00:44:57,680 --> 00:44:58,600
It is not the whole book.

1286
00:44:58,600 --> 00:45:00,040
Microsoft is moving toward a world

1287
00:45:00,040 --> 00:45:01,440
where general frontier capabilities

1288
00:45:01,440 --> 00:45:03,520
sits right alongside the specialized layer.

1289
00:45:03,520 --> 00:45:05,080
It is not an either/or strategy.

1290
00:45:05,080 --> 00:45:07,680
We see the same expansion happening on the runtime side.

1291
00:45:07,680 --> 00:45:10,800
5.4 is not staying a small line-up of two or three models.

1292
00:45:10,800 --> 00:45:13,440
It is already growing toward 10 different versions.

1293
00:45:13,440 --> 00:45:17,080
These will span from 3.8 billion parameters up to 15 billion.

1294
00:45:17,080 --> 00:45:20,080
And every single one of them is still MIT licensed.

1295
00:45:20,080 --> 00:45:23,240
That licensing detail is just as important now as it was earlier.

1296
00:45:23,240 --> 00:45:26,560
It means this isn't just Microsoft growing its own internal layer.

1297
00:45:26,560 --> 00:45:28,680
It is a runtime layer that any developer can use.

1298
00:45:28,680 --> 00:45:30,280
You can pick the size that fits your hardware

1299
00:45:30,280 --> 00:45:32,560
without needing a legal team to sign off first.

1300
00:45:32,560 --> 00:45:35,160
But here is a signal you should pay close attention to.

1301
00:45:35,160 --> 00:45:36,680
Look at the Mayo Clinic partnership.

1302
00:45:36,680 --> 00:45:38,680
Microsoft is not just licensing a generic model

1303
00:45:38,680 --> 00:45:40,360
to a hospital and walking away.

1304
00:45:40,360 --> 00:45:42,840
They are jointly building a frontier model specifically

1305
00:45:42,840 --> 00:45:43,800
for healthcare.

1306
00:45:43,800 --> 00:45:45,520
They are combining frontier reasoning

1307
00:45:45,520 --> 00:45:47,120
with Mayo's own clinical data

1308
00:45:47,120 --> 00:45:49,120
and deploying it at a massive hospital scale.

1309
00:45:49,120 --> 00:45:49,920
This is not a demo.

1310
00:45:49,920 --> 00:45:51,520
This is the reason and runtime pattern

1311
00:45:51,520 --> 00:45:54,560
applied to one of the highest stakes domains in the world.

1312
00:45:54,560 --> 00:45:57,360
You have a major institution putting its name on the outcome.

1313
00:45:57,360 --> 00:45:58,840
So what does that tell us about the future?

1314
00:45:58,840 --> 00:46:02,480
It points toward vertical specific pairs of reason and runtime.

1315
00:46:02,480 --> 00:46:04,800
We are going to see the show up industry by industry.

1316
00:46:04,800 --> 00:46:06,240
Healthcare gets its own tuned pairing

1317
00:46:06,240 --> 00:46:07,320
built on clinical data.

1318
00:46:07,320 --> 00:46:08,520
Legal gets its own.

1319
00:46:08,520 --> 00:46:11,160
Manufacturing, logistics and financial services

1320
00:46:11,160 --> 00:46:13,040
will likely follow the same path.

1321
00:46:13,040 --> 00:46:15,760
Each industry ends up with a frontier reasoning layer

1322
00:46:15,760 --> 00:46:17,800
shaped around its specific problems.

1323
00:46:17,800 --> 00:46:20,480
And that layer is paired with a fast local runtime

1324
00:46:20,480 --> 00:46:22,080
tuned for daily execution.

1325
00:46:22,080 --> 00:46:23,720
If you push that idea forward,

1326
00:46:23,720 --> 00:46:25,960
you can see what it looks like inside a single company.

1327
00:46:25,960 --> 00:46:27,560
You won't have one shared model bolted

1328
00:46:27,560 --> 00:46:30,200
onto every department just because that is what the license allowed.

1329
00:46:30,200 --> 00:46:32,800
Instead, every department runs its own tuned runtime.

1330
00:46:32,800 --> 00:46:35,040
It is shaped around their specific workflows,

1331
00:46:35,040 --> 00:46:37,280
just like the Excel and McKinsey examples.

1332
00:46:37,280 --> 00:46:39,040
Each of those runtimes connects back

1333
00:46:39,040 --> 00:46:41,480
to a shared reasoning layer underneath.

1334
00:46:41,480 --> 00:46:44,480
That foundation provides the deep planning capability

1335
00:46:44,480 --> 00:46:48,000
that the individual runtimes don't need to carry themselves.

1336
00:46:48,000 --> 00:46:49,920
Finance gets a runtime for finance.

1337
00:46:49,920 --> 00:46:51,440
Support gets one for support.

1338
00:46:51,440 --> 00:46:52,480
Underneath all of them,

1339
00:46:52,480 --> 00:46:54,480
the whole company draws from one reasoning foundation

1340
00:46:54,480 --> 00:46:55,960
instead of rebuilding it every time.

1341
00:46:55,960 --> 00:46:57,320
That is the direction of travel.

1342
00:46:57,320 --> 00:46:59,600
It is not one giant model absorbing everything.

1343
00:46:59,600 --> 00:47:02,320
It is also not a scattered pile of small disconnected tools.

1344
00:47:02,320 --> 00:47:03,120
It is a structure.

1345
00:47:03,120 --> 00:47:04,920
You have many tuned runtimes on top

1346
00:47:04,920 --> 00:47:06,840
and one shared reasoning layer underneath.

1347
00:47:06,840 --> 00:47:09,480
It happens industry by industry and department by department.

1348
00:47:09,480 --> 00:47:10,680
Now we need to step back.

1349
00:47:10,680 --> 00:47:13,120
We need to look at how the models, the routing,

1350
00:47:13,120 --> 00:47:15,920
the governance and the trust gap all fit together.

1351
00:47:15,920 --> 00:47:17,160
They aren't just parts.

1352
00:47:17,160 --> 00:47:19,120
They are one single architecture.

1353
00:47:19,120 --> 00:47:20,680
Tying the architecture together.

1354
00:47:20,680 --> 00:47:22,920
Here is the whole stack laid out in one place.

1355
00:47:22,920 --> 00:47:24,960
You have models that split by design.

1356
00:47:24,960 --> 00:47:28,560
You have a runtime layer built to execute as close to the work as possible.

1357
00:47:28,560 --> 00:47:30,760
There is an orchestration layer that decides

1358
00:47:30,760 --> 00:47:32,640
where each request needs to go.

1359
00:47:32,640 --> 00:47:36,000
You have a governance layer trying to solve sovereignty questions and trust gaps.

1360
00:47:36,000 --> 00:47:39,480
And at the very bottom, you have silicon built in lockstep with the models.

1361
00:47:39,480 --> 00:47:42,760
The models aren't trying to fit into whatever hardware was available.

1362
00:47:42,760 --> 00:47:44,000
The hardware was built for them.

1363
00:47:44,000 --> 00:47:46,840
These are five layers and none of them work in isolation.

1364
00:47:46,840 --> 00:47:48,600
If you skipped to this part of the video,

1365
00:47:48,600 --> 00:47:50,560
here is the one sentence you need to hear.

1366
00:47:50,560 --> 00:47:53,000
Five four executes and MI1 decides.

1367
00:47:53,000 --> 00:47:54,120
That is the entire reframe.

1368
00:47:54,120 --> 00:47:55,280
They are not competitors.

1369
00:47:55,280 --> 00:47:57,120
This is not a contest of big versus small.

1370
00:47:57,120 --> 00:47:59,840
It is not a leaderboard with a winner and a loser.

1371
00:47:59,840 --> 00:48:02,760
These are two different jobs based on two different philosophies.

1372
00:48:02,760 --> 00:48:04,280
They are working together on purpose,

1373
00:48:04,280 --> 00:48:06,320
but we should be very clear about something before we finish.

1374
00:48:06,320 --> 00:48:07,680
This is not unique to Microsoft.

1375
00:48:07,680 --> 00:48:10,920
Every major AI provider is moving towards some version of this same split.

1376
00:48:10,920 --> 00:48:13,520
You have a heavy reasoning layer for the hard problems

1377
00:48:13,520 --> 00:48:16,040
and a fast specialized layer for everything else.

1378
00:48:16,040 --> 00:48:18,880
The economics of AI force everyone toward this answer eventually.

1379
00:48:18,880 --> 00:48:21,680
Microsoft just happened to turn it into a product first.

1380
00:48:21,680 --> 00:48:24,400
They did it at their own scale and on their own chips.

1381
00:48:24,400 --> 00:48:26,400
The pattern belongs to the whole industry,

1382
00:48:26,400 --> 00:48:27,800
not just one company road map.

1383
00:48:27,800 --> 00:48:30,880
Because of that, the competitive question has changed.

1384
00:48:30,880 --> 00:48:33,680
It is no longer about who has access to the models.

1385
00:48:33,680 --> 00:48:35,320
The organizations that win the next few years

1386
00:48:35,320 --> 00:48:37,040
won't be the ones with the best license.

1387
00:48:37,040 --> 00:48:39,000
Model access is basically a commodity now.

1388
00:48:39,000 --> 00:48:40,920
If you have a budget, you have a model.

1389
00:48:40,920 --> 00:48:42,840
What is not a commodity is your rooting logic.

1390
00:48:42,840 --> 00:48:44,680
The real value is in the classification work

1391
00:48:44,680 --> 00:48:45,960
and the governance decisions.

1392
00:48:45,960 --> 00:48:47,440
It is the willingness to build the layer

1393
00:48:47,440 --> 00:48:49,400
that actually decides where a request goes.

1394
00:48:49,400 --> 00:48:50,840
That part does not come prebuilt.

1395
00:48:50,840 --> 00:48:53,480
That is what separates companies that are structurally faster

1396
00:48:53,480 --> 00:48:55,360
from companies that still send every request

1397
00:48:55,360 --> 00:48:56,880
through one expensive door.

1398
00:48:56,880 --> 00:48:58,600
Now I want to add a quick caveat here.

1399
00:48:58,600 --> 00:49:01,200
This deserves honesty rather than just acting like everything

1400
00:49:01,200 --> 00:49:02,080
is a settled fact.

1401
00:49:02,080 --> 00:49:03,800
Some of what we talked about is reported.

1402
00:49:03,800 --> 00:49:05,280
It comes from post-built coverage,

1403
00:49:05,280 --> 00:49:08,040
analyst reports and public statements from Microsoft.

1404
00:49:08,040 --> 00:49:11,160
It is not all confirmed line by line in technical docs yet.

1405
00:49:11,160 --> 00:49:13,040
The 5.4 architecture is well documented,

1406
00:49:13,040 --> 00:49:14,800
but some of the newer MEI details

1407
00:49:14,800 --> 00:49:16,440
and the exact benchmark numbers

1408
00:49:16,440 --> 00:49:18,760
are still coming out as the system matures.

1409
00:49:18,760 --> 00:49:20,000
It is worth watching,

1410
00:49:20,000 --> 00:49:21,600
but don't treat it as final truth

1411
00:49:21,600 --> 00:49:24,000
just because the numbers sound precise.

1412
00:49:24,000 --> 00:49:25,360
That distinction is important,

1413
00:49:25,360 --> 00:49:27,040
but it does not change the core point.

1414
00:49:27,040 --> 00:49:29,560
The architecture pattern is real, it stays real,

1415
00:49:29,560 --> 00:49:31,800
regardless of what the final numbers look like.

1416
00:49:31,800 --> 00:49:34,240
And that leaves one big question on the table.

1417
00:49:34,240 --> 00:49:36,080
If model access is a solved problem

1418
00:49:36,080 --> 00:49:39,000
and rooting is the real battleground, what does that change?

1419
00:49:39,000 --> 00:49:40,800
It changes how organizations have to think.

1420
00:49:40,800 --> 00:49:42,240
It isn't just about AI anymore,

1421
00:49:42,240 --> 00:49:45,480
it is about how you build anything from here on out.

1422
00:49:45,480 --> 00:49:48,440
The real shift, nobody's naming, zoom all the way out.

1423
00:49:48,440 --> 00:49:51,720
There is a bigger shift buried underneath everything we just covered.

1424
00:49:51,720 --> 00:49:53,240
And it is not really about Microsoft,

1425
00:49:53,240 --> 00:49:56,440
this is the end of the pick one AI vendor era.

1426
00:49:56,440 --> 00:49:58,160
That instinct made sense a few years ago

1427
00:49:58,160 --> 00:49:59,840
when you were just picking a horse.

1428
00:49:59,840 --> 00:50:02,520
One contract, one model, one API key and you were done.

1429
00:50:02,520 --> 00:50:03,640
That world is over.

1430
00:50:03,640 --> 00:50:06,040
It is over because the winning architecture

1431
00:50:06,040 --> 00:50:07,680
was never going to be a single model.

1432
00:50:07,680 --> 00:50:09,080
It was always going to be a system,

1433
00:50:09,080 --> 00:50:10,440
layers doing different jobs.

1434
00:50:10,440 --> 00:50:13,360
And the system does not come from a single vendor relationship.

1435
00:50:13,360 --> 00:50:16,920
It comes from design, which points to the actual shift worth naming.

1436
00:50:16,920 --> 00:50:19,200
Architectural literacy is becoming a business skill,

1437
00:50:19,200 --> 00:50:20,600
not just a technical one.

1438
00:50:20,600 --> 00:50:25,560
For years, understanding AI meant knowing which model scored higher on a benchmark.

1439
00:50:25,560 --> 00:50:28,160
That conversation lived entirely inside engineering,

1440
00:50:28,160 --> 00:50:30,600
but that is not where the conversation lives anymore.

1441
00:50:30,600 --> 00:50:34,400
Knowing why a request should go to a fast local model instead of a frontier one,

1442
00:50:34,400 --> 00:50:36,600
knowing what rooting logic costs to skip,

1443
00:50:36,600 --> 00:50:39,760
knowing where governance sits relative to where reasoning happens.

1444
00:50:39,760 --> 00:50:41,120
That is business judgment now.

1445
00:50:41,120 --> 00:50:43,080
It shows up in budget meetings and board decks.

1446
00:50:43,080 --> 00:50:44,920
The people who need to understand this split

1447
00:50:44,920 --> 00:50:47,000
are no longer just the ones writing the code.

1448
00:50:47,000 --> 00:50:49,120
Go back to how this whole episode opened.

1449
00:50:49,120 --> 00:50:51,840
The question everyone kept asking was which model wins?

1450
00:50:51,840 --> 00:50:54,560
M-A-I-1 or F-I-4?

1451
00:50:54,560 --> 00:50:55,880
Bigger or smaller?

1452
00:50:55,880 --> 00:50:57,320
Faster or smarter?

1453
00:50:57,320 --> 00:50:59,800
That question was never the real one.

1454
00:50:59,800 --> 00:51:02,480
The real question, the one worth asking from the very first minute,

1455
00:51:02,480 --> 00:51:04,760
was how these systems divide labor.

1456
00:51:04,760 --> 00:51:05,720
It is not a contest.

1457
00:51:05,720 --> 00:51:07,000
It is a design decision.

1458
00:51:07,000 --> 00:51:08,920
Every section since has just been unpacking

1459
00:51:08,920 --> 00:51:10,960
what that division looks like in practice.

1460
00:51:10,960 --> 00:51:14,040
What it costs to get wrong and what it is worth to get right.

1461
00:51:14,040 --> 00:51:16,960
So here is the belief breaking statement worth sitting with.

1462
00:51:16,960 --> 00:51:19,440
The companies still asking which model is smarter

1463
00:51:19,440 --> 00:51:22,360
are already behind the ones asking how to root intelligently.

1464
00:51:22,360 --> 00:51:23,720
That gap is not going to close.

1465
00:51:23,720 --> 00:51:26,680
It is going to widen quietly, one quarter at a time.

1466
00:51:26,680 --> 00:51:29,520
Eventually the company is still fighting over leaderboard positions

1467
00:51:29,520 --> 00:51:32,680
will look up and realize the race was never about the leaderboard.

1468
00:51:32,680 --> 00:51:34,280
So here is what to actually do with this.

1469
00:51:34,280 --> 00:51:37,200
Take your current A-I workloads and sort them into two piles.

1470
00:51:37,200 --> 00:51:40,440
Needs reasoning, needs fast execution.

1471
00:51:40,440 --> 00:51:42,480
Most of what is running through a frontier model

1472
00:51:42,480 --> 00:51:44,280
right now belongs in that second pile.

1473
00:51:44,280 --> 00:51:47,280
Pick one high volume low complexity task this month.

1474
00:51:47,280 --> 00:51:49,960
Test it against a small model before you default

1475
00:51:49,960 --> 00:51:51,800
to the expensive option out of habit.

1476
00:51:51,800 --> 00:51:54,040
If this changed how you think about your own A-I stack,

1477
00:51:54,040 --> 00:51:56,120
follow me, Möko Peters on LinkedIn.

1478
00:51:56,120 --> 00:51:58,320
And if you want more of this, leave a review.

1479
00:51:58,320 --> 00:51:59,600
It helps more people find it.

1480
00:51:59,600 --> 00:52:01,320
The architecture is already being built.

1481
00:52:01,320 --> 00:52:04,520
The only open question left is who builds the routing layer first?

