1
00:00:00,000 --> 00:00:04,240
Hello everyone and welcome to another episode of Microsoft Knowledge Nuggets here on M365.

2
00:00:04,240 --> 00:00:08,160
FM, I'm Mirko Peters. Imagine you need one virtual machine for a small app.

3
00:00:08,160 --> 00:00:11,680
You pick a region, choose a VM size, see the price and create it.

4
00:00:11,680 --> 00:00:16,240
Simple, right? Now imagine that same app suddenly needs 2,000 VMs overnight.

5
00:00:16,240 --> 00:00:20,000
You might be running simulations, processing a huge data set or testing software with

6
00:00:20,000 --> 00:00:24,640
thousands of short-lived workers. Suddenly that one easy VM choice turns into a real headache.

7
00:00:24,640 --> 00:00:27,360
The size you picked might not have enough capacity in your region.

8
00:00:27,360 --> 00:00:31,680
A different size could work, but the price shifts. Spot VMs cost way less, but Azure can take

9
00:00:31,680 --> 00:00:36,080
them back at any time. You could build multiple virtual machine scale sets and write custom scripts

10
00:00:36,080 --> 00:00:40,560
to hunt for capacity, but that quickly becomes a management nightmare. Here's the question,

11
00:00:40,560 --> 00:00:45,440
Azure Compute Fleet is built to answer. Why should one exact VM size keep a large deployment

12
00:00:45,440 --> 00:00:49,360
from starting at all? By the end of this episode, you'll know what Compute Fleet actually does,

13
00:00:49,360 --> 00:00:52,640
what it doesn't do, and where it fits alongside virtual machine scale sets.

14
00:00:53,280 --> 00:00:58,800
Compute Fleet in plain English. Azure Compute Fleet is a managed service designed to get

15
00:00:58,800 --> 00:01:03,120
you a large pool of VM compute capacity. Notice I didn't say it creates one specific machine.

16
00:01:03,120 --> 00:01:07,200
You tell Azure what kind of capacity you can use, how much you need, where you're willing to run it,

17
00:01:07,200 --> 00:01:11,280
and whether you want standard on-demand VMs, lower cost spot VMs, or both.

18
00:01:11,280 --> 00:01:16,160
Then Azure scans for capacity across those acceptable options. Think of it like a modern office

19
00:01:16,160 --> 00:01:22,560
building. You don't walk in demanding desk 27 on floor 8, in one exact room facing one exact window.

20
00:01:22,560 --> 00:01:27,200
Instead, you tell the building manager, "I need 500 desks for my team." These floors work,

21
00:01:27,200 --> 00:01:31,760
these room layouts are fine. Here's our budget. The manager looks at what's free and gets everyone

22
00:01:31,760 --> 00:01:36,400
seated. Compute Fleet works the same way. Your request might say, "I need a thousand workers,

23
00:01:36,400 --> 00:01:41,040
they need this much CPU and memory. These VM sizes are acceptable. These availability zones are

24
00:01:41,040 --> 00:01:44,800
acceptable. Use on-demand for part of the workload and spot where it makes sense."

25
00:01:44,800 --> 00:01:50,320
That wider net makes it much easier to find available capacity. A VM SKU is just a particular size

26
00:01:50,320 --> 00:01:54,960
and type a virtual machine. It describes things like CPU, memory, local storage, and processor

27
00:01:54,960 --> 00:02:00,000
family. If you ask for only one SKU, Azure can only search for that one exact shape. If that capacity

28
00:02:00,000 --> 00:02:05,920
is busy, your deployment waits or fails. With Compute Fleet, you list multiple acceptable VM sizes.

29
00:02:05,920 --> 00:02:10,560
Azure can then place some workers on one size and others on another as long as your workload runs

30
00:02:10,560 --> 00:02:16,160
on both. For big workloads, that flexibility is the whole point. Compute Fleet can handle up to 10,000

31
00:02:16,160 --> 00:02:20,880
VMs in a single deployment. You no longer need to create thousands of machines one by one or build

32
00:02:20,880 --> 00:02:25,760
separate scripts for every acceptable VM size. Instead, you create one fleet request. Behind the scenes,

33
00:02:25,760 --> 00:02:30,560
Azure checks the capacity choices you allowed, finds a workable mix, and creates the VM groups

34
00:02:30,560 --> 00:02:34,800
that make up the fleet. You still have to provide the VM image, the network setup, and the software

35
00:02:34,800 --> 00:02:39,120
each worker should run. Fleet doesn't understand your business job on its own. It doesn't know whether

36
00:02:39,120 --> 00:02:43,840
a worker is calculating financial risk, rendering a video frame, testing software, or processing a

37
00:02:43,840 --> 00:02:49,120
genome sample. It also doesn't root web traffic, replace a load balancer, or fix code inside a VM.

38
00:02:49,120 --> 00:02:54,400
Fleet's role is narrower. It helps you acquire and manage a large pool of compute capacity when

39
00:02:54,400 --> 00:02:58,800
more than one path can get you there. That's the shift. Instead of handing Azure a rigid shopping

40
00:02:58,800 --> 00:03:04,080
list that says, "I need 1,000 of this exact machine." You give it a practical request that says,

41
00:03:04,080 --> 00:03:08,960
"I need this much compute, and here are the choices that work." The next part is where that

42
00:03:08,960 --> 00:03:13,760
request becomes useful, because you need to decide how much freedom you're willing to give Azure.

43
00:03:13,760 --> 00:03:18,320
The three choices that give Azure room to work. The first choice is the list of VM sizes you'll

44
00:03:18,320 --> 00:03:23,680
accept. Most people start with a name like D8's, F16, or some exact SKU that makes sense when you're

45
00:03:23,680 --> 00:03:28,000
running one server and you know exactly what it needs. But at Fleet scale, the better question

46
00:03:28,000 --> 00:03:33,680
isn't "which processor name do I want?" It's, "What does my work actually need?" Imagine a risk

47
00:03:33,680 --> 00:03:39,120
model that runs thousands of separate calculations. Each worker needs enough, CPU and memory to process

48
00:03:39,120 --> 00:03:43,520
one slice of the data. It might run just as well on several compute-focused VM families,

49
00:03:43,520 --> 00:03:48,400
even if the underlying processor is different. The job cares about finishing calculations correctly.

50
00:03:48,400 --> 00:03:52,720
It doesn't care whether every worker has the same label on the box. So instead of allowing one VM

51
00:03:52,720 --> 00:03:56,800
size, you can allow several sizes that meet the same practical need that gives as you

52
00:03:56,800 --> 00:04:00,720
have more capacity options to search through, while your workers still get the resources they need.

53
00:04:00,720 --> 00:04:06,160
There is a limit to this flexibility though. Your software has to run correctly on every VM type you allow.

54
00:04:06,160 --> 00:04:09,680
If one family needs a special driver, a different processor design or a particular

55
00:04:09,680 --> 00:04:13,840
local disk setup, don't add it just because it looks cheaper. A wider list only helps when every

56
00:04:13,840 --> 00:04:18,240
option you include is a real option. The second choice is what you want Azure to favor when it finds

57
00:04:18,240 --> 00:04:22,880
those options. Sometimes finishing the work matters more than finding the lowest possible price.

58
00:04:22,880 --> 00:04:26,720
Maybe a team meets to complete an overnight calculation before the market opens. In that case,

59
00:04:26,720 --> 00:04:31,680
you'd lean toward a capacity-focused choice. Your telling Azure find available compute that meets my

60
00:04:31,680 --> 00:04:37,520
rules. The job needs to start. This doesn't mean Azure creates capacity from nowhere. It means Azure

61
00:04:37,520 --> 00:04:42,080
can search through the sizes, zones and purchase choices you allowed, then favor the options with the

62
00:04:42,080 --> 00:04:46,720
best chance of getting your workers running. Other jobs can take a different path. Maybe you're creating

63
00:04:46,720 --> 00:04:51,840
test environments, running a large simulation that can retry later, or processing a backlog of

64
00:04:51,840 --> 00:04:57,760
files with no hard deadline. Cost may matter more than speed in those cases. A price-focused choice

65
00:04:57,760 --> 00:05:02,880
tells Azure to favor lower cost eligible capacity. That can reduce your compute bill, but it can

66
00:05:02,880 --> 00:05:07,200
also mean waiting longer for enough workers. The least expensive choice isn't always available when

67
00:05:07,200 --> 00:05:11,440
you need it, and it may not give you the full amount of capacity immediately. You've probably seen

68
00:05:11,440 --> 00:05:15,840
this outside Azure. A cheap flight is perfect if your schedule is flexible. If you need to be

69
00:05:15,840 --> 00:05:20,160
somewhere for an important meeting tomorrow morning, the lowest fare probably isn't the decision

70
00:05:20,160 --> 00:05:24,320
that helps you. Compute fleet lets you make the same trade-off with compute. Then there's the middle

71
00:05:24,320 --> 00:05:29,200
ground. Many teams don't want to pay the highest price for every worker, but they also can't accept a

72
00:05:29,200 --> 00:05:34,160
fleet that never reaches a useful size. A balanced approach gives Azure a room to consider both price

73
00:05:34,160 --> 00:05:38,320
and available capacity. You set the boundaries and Azure works inside them. That's a much more

74
00:05:38,320 --> 00:05:42,720
honest design than pretending every workload has one simple goal. Some jobs need a deadline,

75
00:05:42,720 --> 00:05:48,080
some jobs need a budget, most need both. The third choice is how you describe the machines themselves.

76
00:05:48,080 --> 00:05:52,160
You can list exact SKUs, which works well when you've already tested a small set of sizes,

77
00:05:52,160 --> 00:05:55,840
but Azure also has an attribute-based approach, currently available in Preview,

78
00:05:55,840 --> 00:05:59,920
where you describe the shape of compute you need. For example, you might describe a worker

79
00:05:59,920 --> 00:06:05,520
that needs a certain amount of CPU, memory, local storage, or an accelerator like a GPU. Azure can

80
00:06:05,520 --> 00:06:10,640
then find VM sizes that match those requirements. Think about the difference. An SKU list says,

81
00:06:10,640 --> 00:06:15,920
use these named models. An attribute request says, "Find machines with these physical needs."

82
00:06:15,920 --> 00:06:20,240
The second approach can reduce the work of updating your fleet definition when newer VM

83
00:06:20,240 --> 00:06:24,400
generations become available, but you still need to test your image and software against the kinds

84
00:06:24,400 --> 00:06:28,560
of machines as you may choose. Don't confuse flexibility with handing over every decision.

85
00:06:28,560 --> 00:06:33,200
You define the safe boundaries and Azure selects within them. So those are the three choices,

86
00:06:33,200 --> 00:06:37,120
which VM sizes or shapes will work, whether capacity or price comes first,

87
00:06:37,120 --> 00:06:42,000
and how broadly you describe the compute you need. Lower cost capacity can make a fleet far cheaper,

88
00:06:42,000 --> 00:06:46,720
but it comes with one rule that every team needs to plan for. Sometimes Azure can take a worker away.

89
00:06:46,720 --> 00:06:51,440
Spot and on demand, cheap seats and reserved seats. So let's talk about the two ways you can

90
00:06:51,440 --> 00:06:55,920
pay for the VMs in a fleet. On demand VMs are the stable option. You ask Azure for a VM,

91
00:06:55,920 --> 00:07:00,080
you pay the normal rate while it runs, and Azure doesn't reclaim it just because another customer needs

92
00:07:00,080 --> 00:07:05,040
that capacity. For work that must keep moving, on demand capacity gives you a dependable base.

93
00:07:05,040 --> 00:07:08,960
Think of it like booking confirmed seats for the people who must be at the event. Your job

94
00:07:08,960 --> 00:07:13,120
controller might run there, your queue service might depend on it, or maybe you need a minimum number

95
00:07:13,120 --> 00:07:16,960
of workers running all night because missing the morning deadline isn't an option. Spot VMs work

96
00:07:16,960 --> 00:07:22,320
differently. Azure sells unused compute capacity at a lower price, sometimes much lower than the normal

97
00:07:22,320 --> 00:07:26,640
on demand rate. That makes Spot useful when you have a lot of work to process and can handle

98
00:07:26,640 --> 00:07:31,360
interruptions. But the lower price comes with a condition. Azure can reclaim a Spot VM when it

99
00:07:31,360 --> 00:07:35,760
needs that capacity back. The VM can disappear while it's working. That doesn't mean Spot is

100
00:07:35,760 --> 00:07:40,560
unreliable in a bad way. It means you need to use it honestly. A Spot VM is not the place for the only

101
00:07:40,560 --> 00:07:45,600
copy of an important database, a single user facing server, or a job that loses hours of progress

102
00:07:45,600 --> 00:07:50,480
when one worker stops. You plan for the worker to leave. This is why a mixed fleet often makes sense.

103
00:07:50,480 --> 00:07:55,360
You keep enough on demand VMs for the work that must continue. Then you add Spot VMs for the

104
00:07:55,360 --> 00:07:59,840
extra processing power. When Spot capacity is available, the fleet can process much more work for

105
00:07:59,840 --> 00:08:05,120
less money. When some Spot workers disappear, the baseline on demand workers keep going. That setup

106
00:08:05,120 --> 00:08:09,680
gives you a practical balance. The stable workers protect your minimum throughput while the Spot

107
00:08:09,680 --> 00:08:13,760
workers help you finish sooner or process more work when Azure has spare capacity.

108
00:08:13,760 --> 00:08:18,560
Compute Fleet helps manage the capacity side of that mix. It can use the allocation choices you

109
00:08:18,560 --> 00:08:23,440
set, look for eligible Spot or on-demand options, and react when Spot VMs are evicted.

110
00:08:23,440 --> 00:08:28,480
Still, Fleet cannot protect work that exists only inside of VM. Your workload design has to do that.

111
00:08:28,480 --> 00:08:32,160
Imagine a video rendering job. One worker renders frame one through frame one hundred.

112
00:08:32,160 --> 00:08:35,920
Another handles the next group. If a Spot worker disappears halfway through a frame,

113
00:08:35,920 --> 00:08:39,200
the system should mark that frame as unfinished and send it back to the queue.

114
00:08:39,200 --> 00:08:44,480
The next worker starts it again. No drama, no loss project. That same pattern works for simulations,

115
00:08:44,480 --> 00:08:49,040
data extraction and transformation jobs, software builds, automated tests, development

116
00:08:49,040 --> 00:08:53,440
environments and genomics processing where many samples can run separately. These jobs work

117
00:08:53,440 --> 00:08:58,000
well because they can split into smaller pieces. One worker can process one file, one worker can

118
00:08:58,000 --> 00:09:02,480
calculate one scenario, one worker can test one build if the worker goes away, another worker can

119
00:09:02,480 --> 00:09:06,880
pick up that small piece later. But a single database server is a very different job.

120
00:09:06,880 --> 00:09:11,280
If the database stores its only copy of the data on a Spot VM, an eviction can interrupt the

121
00:09:11,280 --> 00:09:15,600
service and put recovery at risk. The same applies to stateful software that cannot restart

122
00:09:15,600 --> 00:09:21,120
cleanly or a customer facing application that depends on one VM staying online. Spot doesn't

123
00:09:21,120 --> 00:09:26,000
turn those designs into safe designs. For a Spot-friendly workload, keep progress somewhere durable

124
00:09:26,000 --> 00:09:30,160
outside the worker VM. That might mean a storage account, a database or another service that

125
00:09:30,160 --> 00:09:35,440
records what has finished and what still needs work. Use a durable queue so workers pull tasks one at a time.

126
00:09:35,440 --> 00:09:40,960
Save checkpoints during long running jobs. Make each task safe to retry, so running it again doesn't

127
00:09:40,960 --> 00:09:46,320
create duplicate results or corrupt data. That way, when Azure reclaims a Spot VM, you lose only

128
00:09:46,320 --> 00:09:50,480
the small piece it was processing, not the entire job. This can save a lot of money but cost alone

129
00:09:50,480 --> 00:09:54,560
doesn't tell you whether compute fleet belongs in your design. The next question is whether you need

130
00:09:54,560 --> 00:09:59,840
to fleet at all or whether a virtual machine scale set fits the job better. Compute fleet versus

131
00:09:59,840 --> 00:10:04,960
virtual machine scale sets. You might be wondering, doesn't Azure already have a service called virtual

132
00:10:04,960 --> 00:10:09,840
machine scale sets? It does and the two are related because they both help you run groups of Azure VMs

133
00:10:09,840 --> 00:10:14,560
instead of managing each one by hand. But here's the thing, they solve different problems starting from

134
00:10:14,560 --> 00:10:19,600
different questions. A virtual machine scale set or VMSS for short starts with your application.

135
00:10:19,600 --> 00:10:24,800
Let's say you have a web app, an API or an internal business service, you create a group of VMs with

136
00:10:24,800 --> 00:10:30,000
a similar setup, put them behind a load balancer and let the group grow or shrink as demand changes.

137
00:10:30,000 --> 00:10:34,720
Traffic rises, so the scale set adds more application servers, traffic falls and it removes servers,

138
00:10:34,720 --> 00:10:38,880
so you aren't paying for machines you no longer need. That's a very normal pattern for a long running

139
00:10:38,880 --> 00:10:43,760
application tier. Your VMs usually run the same image, the same application and a closely related

140
00:10:43,760 --> 00:10:48,720
VM setup. Azure can use health checks to see whether the app is responding and it can add or

141
00:10:48,720 --> 00:10:54,080
remove instances based on CPU use, a schedule or another metric you choose. Think of a VM scale set

142
00:10:54,080 --> 00:10:58,640
like a store manager planning a regular team. Every morning the manager needs cashiers, shelf

143
00:10:58,640 --> 00:11:02,960
stockers and people at the service desk. They know the jobs, the location and the basic tools each

144
00:11:02,960 --> 00:11:07,440
person needs. When the store gets busy, they call in more people from the same team. That's what VM

145
00:11:07,440 --> 00:11:12,480
assess does well. It runs a known application pattern and adjusts the number of workers as demand

146
00:11:12,480 --> 00:11:16,880
changes. Compute fleets start somewhere else. Instead of asking out how many web servers does my app

147
00:11:16,880 --> 00:11:22,160
need right now? It asks how can I get a very large amount of acceptable compute capacity for this job?

148
00:11:22,160 --> 00:11:26,960
That distinction changes how you design the request. With a fleet the VM size can vary across the

149
00:11:26,960 --> 00:11:31,600
workers as long as each worker can perform the task. You might accept several machine families,

150
00:11:31,600 --> 00:11:36,160
use different purchase options and let Azure choose a workable mix based on your rules.

151
00:11:36,160 --> 00:11:40,400
Think of compute fleet as a staffing agency trying to fill thousands of shifts before a deadline.

152
00:11:40,400 --> 00:11:44,720
The agency doesn't care whether every worker wears the same shoe size, it cares whether each

153
00:11:44,720 --> 00:11:49,200
person can do the work, whether enough people arrive on time and what the full shift costs.

154
00:11:49,200 --> 00:11:53,840
That's a much better fit for a job that can split across many independent workers, a large data run,

155
00:11:53,840 --> 00:11:58,320
a simulation, a rendering project or a big set of automated tests. The difference isn't that one

156
00:11:58,320 --> 00:12:02,400
service is better than the other, they solve different problems. So choose virtual machine scale

157
00:12:02,400 --> 00:12:06,000
sets when you run a steady application tier that receives traffic. That could be a customer

158
00:12:06,000 --> 00:12:10,640
facing website, an API, a line of business app or a group of application servers behind Azure

159
00:12:10,640 --> 00:12:15,280
load balancer or application gateway. In those cases you usually care about application health

160
00:12:15,280 --> 00:12:19,520
and traffic handling. You need instances with a known setup, you need a way to add servers when

161
00:12:19,520 --> 00:12:24,880
users arrive and remove them when demand drops. VMSS fits that model. Choose compute fleet when the

162
00:12:24,880 --> 00:12:29,040
main challenge is getting a lot of compute for work that can spread out. The work might last

163
00:12:29,040 --> 00:12:34,000
for hours, days or only a short burst but it doesn't depend on one identical application server

164
00:12:34,000 --> 00:12:38,640
receiving user requests. Fleet works well when a worker can take a task, complete it,

165
00:12:38,640 --> 00:12:42,720
save the result and move on. That's common in high performance computing, batch processing,

166
00:12:42,720 --> 00:12:47,360
large simulations, data processing and large temporary worker pools. There's also a practical point

167
00:12:47,360 --> 00:12:51,360
that trips people up. Compute fleet does not replace your application design. It doesn't replace

168
00:12:51,360 --> 00:12:56,400
the queue that hands out work, the checkpoints that save progress or a load balancer for a web app.

169
00:12:56,400 --> 00:13:00,240
And it doesn't turn a single stateful server into a fault tolerant system. As you can help acquire

170
00:13:00,240 --> 00:13:04,480
the VMS but your software still needs to know what each VMS should do, where its data lives,

171
00:13:04,480 --> 00:13:09,520
and what should happen if a worker stops halfway through a task. That shared responsibility matters.

172
00:13:09,520 --> 00:13:14,400
As your manages the physical servers and the fleet capacity rules, you design the workload so

173
00:13:14,400 --> 00:13:19,280
it can survive the normal events of cloud computing, including a VM restart, a failed task or

174
00:13:19,280 --> 00:13:23,520
an interrupted worker. So a simple decision guide looks like this. If you're scaling an application

175
00:13:23,520 --> 00:13:27,920
because user traffic changes, start with virtual machine scale sets. If you're gathering a large

176
00:13:27,920 --> 00:13:32,160
amount of flexible compute because a job needs many workers, look at compute fleet. And if your

177
00:13:32,160 --> 00:13:37,280
architecture includes both needs, you can use both services. A VM scale set can run the customer

178
00:13:37,280 --> 00:13:42,000
facing application tier, while compute fleet supplies temporary workers for the heavy background job.

179
00:13:42,000 --> 00:13:45,920
Picture a team that needs thousands of workers overnight to finish a calculation before the next

180
00:13:45,920 --> 00:13:49,600
business date. That's where these choices stop being product names and start changing the outcome.

181
00:13:49,600 --> 00:13:55,840
A fleet example. Finishing a large job without betting on one VM size. Imagine a finance team that

182
00:13:55,840 --> 00:14:00,000
runs risk calculations every night. They need to process a huge amount of market data,

183
00:14:00,000 --> 00:14:04,480
test thousands of possible outcomes and produce results before people arrive the next morning.

184
00:14:04,480 --> 00:14:08,240
The team has one clear goal. Finish the calculations by morning without paying the normal

185
00:14:08,240 --> 00:14:12,560
on-demand rate for every single worker. Their first attempt looks simple. They choose one compute

186
00:14:12,560 --> 00:14:17,600
focused VM size, request a large number of workers, and start the job. Then as you cannot find enough

187
00:14:17,600 --> 00:14:22,080
of that exact size in the region. Nothing is wrong with the VM size or the code. There just

188
00:14:22,080 --> 00:14:26,560
isn't enough capacity for that one exact choice when the team needs it. If their whole design

189
00:14:26,560 --> 00:14:31,520
depends on that one SKU, the job starts late. A late start means fewer results by morning or a rushed

190
00:14:31,520 --> 00:14:36,960
and expensive move to another setup. So the team changes the request. Instead of demanding one VM size,

191
00:14:36,960 --> 00:14:42,480
they identify several CPU focused sizes that their risk software can use. Each option has enough

192
00:14:42,480 --> 00:14:47,520
memory for a task, enough CPU for a calculation, and works with the same VM image. They also allow more

193
00:14:47,520 --> 00:14:52,000
than one availability zone where their data and workload can run. That gives Azure more places to

194
00:14:52,000 --> 00:14:57,120
find workers. Next, the team decides how much capacity must stay available no matter what. They choose

195
00:14:57,120 --> 00:15:02,160
a group of on-demand VMs as their reliable base. Those workers give the job a minimum processing rate

196
00:15:02,160 --> 00:15:07,280
throughout the night. Even if lower cost capacity becomes scarce, the run continues. Above that

197
00:15:07,280 --> 00:15:12,240
base they allow spot VMs. Those spot workers do the extra work. When Azure has spare capacity,

198
00:15:12,240 --> 00:15:16,560
the team can process many more calculations without using on-demand VMs for every task.

199
00:15:16,560 --> 00:15:20,960
The request now says something like this. Give us enough stable workers to guarantee a minimum

200
00:15:20,960 --> 00:15:25,440
result, then add as many lower cost workers as you can from this approved list of VM sizes.

201
00:15:25,440 --> 00:15:29,840
That changes the situation. The team isn't betting the whole overnight run on one size, one zone,

202
00:15:29,840 --> 00:15:34,080
or one price type. But there's another part of the design that matters just as much as the fleet

203
00:15:34,080 --> 00:15:39,600
request. The risk calculations cannot live only inside the worker VMs. Before the fleet starts,

204
00:15:39,600 --> 00:15:44,160
a job system breaks the large calculation into small pieces. One task might calculate the outcome

205
00:15:44,160 --> 00:15:49,760
for one market scenario. Another might handle a different portfolio or time range. Each task enters a queue.

206
00:15:49,760 --> 00:15:54,400
A worker takes one task, processes it, writes the result to durable storage,

207
00:15:54,400 --> 00:15:58,720
and then asks for another task. Progress lives outside the worker. That means a worker can

208
00:15:58,720 --> 00:16:02,960
disappear without taking the whole job with it. Supposer spot VM receives an eviction

209
00:16:02,960 --> 00:16:08,240
notice halfway through a calculation. The VM stops. The unfinished task remains marked as incomplete

210
00:16:08,240 --> 00:16:13,520
in the queue. Another worker, perhaps a new spot VM or one of the on-demand workers, can take that

211
00:16:13,520 --> 00:16:18,560
task and run it again. The team loses one small piece of work, but they don't lose the whole overnight

212
00:16:18,560 --> 00:16:23,200
calculation. And boom, that's the real benefit of a fleet-friendly design. Azure handles finding

213
00:16:23,200 --> 00:16:28,560
and managing the compute capacity. The team's job system handles the work, retries, and saved results.

214
00:16:28,560 --> 00:16:33,200
Neither side can do the other side's job. Compute fleet cannot decide whether rerunning a calculation

215
00:16:33,200 --> 00:16:38,240
creates the wrong financial result. The team has to make tasks safe to retry and prevent duplicate

216
00:16:38,240 --> 00:16:42,240
processing. At the same time, the team doesn't need custom scripts that search through every approved

217
00:16:42,240 --> 00:16:46,960
VM size and keep trying to create workers one group at a time. They define the acceptable choices

218
00:16:46,960 --> 00:16:52,480
and let Azure work within those rules. By morning, the exact number of spot workers may have changed

219
00:16:52,480 --> 00:16:57,760
during the night. That's expected. The stable workers kept processing. Extra spot workers increased

220
00:16:57,760 --> 00:17:02,720
throughput when capacity existed, interrupted work returned to the queue. The final cost stayed

221
00:17:02,720 --> 00:17:06,560
under more control than an all-on-demand run. This isn't a promise that every task will finish

222
00:17:06,560 --> 00:17:11,120
at the lowest price. It's a way to avoid tying the whole job to one narrow capacity choice.

223
00:17:11,120 --> 00:17:15,760
The same pattern appears in life sciences. A research team may need to process a large set of

224
00:17:15,760 --> 00:17:20,240
genome samples. Each sample can move through the same analysis pipeline, but each sample is separate

225
00:17:20,240 --> 00:17:25,440
from the next. The team stores input data and results outside the worker VMs. A worker processes one

226
00:17:25,440 --> 00:17:31,120
sample or one stage of a sample, saves the output and moves on. If that worker stops, another worker

227
00:17:31,120 --> 00:17:35,280
resumes from the last save point. Compute fleet can supply the large, flexible worker pool,

228
00:17:35,280 --> 00:17:39,520
but the pipeline still needs to track what finished, what failed, and what must run again.

229
00:17:39,520 --> 00:17:43,360
Before you put a fleet into a real design though, there are a few decisions that prevent you from

230
00:17:43,360 --> 00:17:48,640
choosing it for the wrong kind of workload. When to use it and what to decide first. What exactly

231
00:17:48,640 --> 00:17:52,880
is a compute fleet workload? It's work that can spread across many independent workers. That's the

232
00:17:52,880 --> 00:17:57,440
starting point. Here's the simple question you need to answer. If one worker stops can another worker

233
00:17:57,440 --> 00:18:02,160
take over without you having to fix the job by hand? If the answer is yes, you probably have a good

234
00:18:02,160 --> 00:18:07,200
fleet workload, batch processing, test workers, analysis jobs, and large file processing often fit

235
00:18:07,200 --> 00:18:11,280
because the work splits into separate units that run in parallel. Start with the workload behavior,

236
00:18:11,280 --> 00:18:16,480
not the VM name. Think of it like a team of freelancers. If one drops out, another picks up the task

237
00:18:16,480 --> 00:18:21,520
without losing progress. Can the work run without keeping its state inside one VM? If it needs state,

238
00:18:21,520 --> 00:18:25,680
can it save progress somewhere durable and restart from that point? Can the same task run again

239
00:18:25,680 --> 00:18:30,640
safely if a worker never reports back? Those answers tell you far more than any pricing page ever will.

240
00:18:30,640 --> 00:18:36,160
Now decide the minimum capacity that must stay available. This is the part of the job you cannot

241
00:18:36,160 --> 00:18:41,920
afford to lose if lower cost capacity becomes unavailable or gets interrupted. Use on-demand VMs for

242
00:18:41,920 --> 00:18:47,040
that baseline then decide how much extra capacity you are willing to add through spot VMs. That number

243
00:18:47,040 --> 00:18:51,680
should come from your deadline, your job backlog, and your budget. Not from a guest that spot will always

244
00:18:51,680 --> 00:18:57,280
be there. After that define your VM choices carefully. You might approve several related VM families

245
00:18:57,280 --> 00:19:01,600
that your software already supports. Or if the preview feature fits your design, you can describe

246
00:19:01,600 --> 00:19:06,480
the CPU, memory, storage, or accelerator needs and let azure pick matching machines keep the list

247
00:19:06,480 --> 00:19:11,680
practical. A narrow list limits the places as you can search. An overly broad list can put your

248
00:19:11,680 --> 00:19:16,400
workload on machines that behave differently from the ones you tested. Every allowed choice needs

249
00:19:16,400 --> 00:19:21,360
to work with your image, your network setup, and your software. You also need a clear ceiling.

250
00:19:21,360 --> 00:19:25,680
Set the target capacity for the job then set a maximum cost exposure before launching.

251
00:19:25,680 --> 00:19:30,720
Large compute jobs can grow fast when a queue contains more work than expected. Or when a deadline

252
00:19:30,720 --> 00:19:35,680
pushes more work onto on-demand capacity. A budget is not just a fine and setting, it is a design

253
00:19:35,680 --> 00:19:40,640
boundary. Where your workload and region supported spread the fleet across availability zones.

254
00:19:40,640 --> 00:19:45,360
This gives azure more placement choices and reduces the chance that one local capacity issue affects

255
00:19:45,360 --> 00:19:50,720
every worker. Keep an eye on the job, not just the VM count. Monitor the capacity fleet has actually

256
00:19:50,720 --> 00:19:56,480
acquired. Watch the queue backlog, failed tasks, retry count, eviction events, completion rate,

257
00:19:56,480 --> 00:20:01,040
and cost. A fleet with many healthy VMs can still be failing the business goal if the queue is growing

258
00:20:01,040 --> 00:20:05,280
or results are not being saved. Here's the mistake most people make. They treat compute fleet as

259
00:20:05,280 --> 00:20:10,160
automatic application recovery. It does not do that. It manages capacity. Your job system still

260
00:20:10,160 --> 00:20:14,400
needs to recover the work. Once those pieces are in place, one flexible request can replace a

261
00:20:14,400 --> 00:20:19,200
pile of manual skew checks, separate spot plans and scripts that retry capacity one VM group at a time.

262
00:20:19,200 --> 00:20:25,520
So pick one workload that already runs as a batch job. Or one that you can safely retry.

263
00:20:25,520 --> 00:20:30,320
Write down five VM choices that can genuinely run it. Not five names you found in a list,

264
00:20:30,320 --> 00:20:35,440
but five tested options your software supports. Then next to that list write down four more things.

265
00:20:35,440 --> 00:20:41,040
The minimum on demand capacity you need. The extra spot capacity you can use. Where progress is saved.

266
00:20:41,040 --> 00:20:46,080
And exactly how an unfinished task returns for retry. Add a cost limit before you launch.

267
00:20:46,080 --> 00:20:50,560
That small design exercise turns compute fleet from a product name into a real plan.

268
00:20:50,560 --> 00:20:56,320
It changes the request from I need this exact VM to I need this work completed and these compute

269
00:20:56,320 --> 00:21:01,440
choices are safe for long running application tiers that serve user traffic continue with virtual

270
00:21:01,440 --> 00:21:06,160
machine scale sets next Azure can find the workers but your workload design decides what happens when

271
00:21:06,160 --> 00:21:06,960
one disappears.

