Scaling and resources
How functions sleep and wake, how many replicas they run, the CPU and memory they get, and how to change those settings.
Settings a new function starts with
| Setting | Value |
|---|---|
| Minimum replicas | 0 |
| Maximum replicas | 10 |
| Cooldown | 300 seconds |
| CPU | 0.25 CPU (250m) |
| Memory | 256 MiB (256Mi) |
These come from the platform, not from the namespace — see Namespace defaults.
Scale to zero
With a minimum of 0 replicas:
- The function sleeps, with the status Sleeping, until something calls it. An HTTP call or a
timer run wakes it: the namespace's proxy starts one replica and holds the call until the
replica is ready, for up to 30 seconds. After that the caller gets
504. - When no call or timer run has reached the function for the length of its cooldown, the proxy stops the replica again. It checks every 30 seconds, so a function can run up to half a minute longer than its cooldown.
- A proxy counts only the calls that pass through itself. With more than one function namespace in the resource group, another namespace's proxy can stop a function that is being called — see How a request reaches a function.
The first call after a sleep takes as long as the function needs to start.
Important
A function built from the portal does not go back to sleep. After its first deploy it sleeps until its first call; from then on it keeps one replica running, and holds that replica's CPU and memory, whatever its cooldown. The platform refreshes the function's status about once a minute and the proxy counts each refresh as activity, so a cooldown longer than about a minute never runs out. A function that runs an image from your own pipeline is not affected: it sleeps after its cooldown as described above.
Keep a function warm
With a minimum above 0, the function always runs that many replicas and never sleeps. That avoids the wait on the first call, at the cost of CPU and memory held all the time.
Maximum replicas
The platform does not add replicas under load. A function runs its minimum number of replicas, or a single one while it is awake when the minimum is 0. The maximum is stored and shown, but it does not make a function scale out.
CPU and memory
A function's CPU and memory are both its request and its limit: it is given exactly that much. A replica that needs more memory than its limit is stopped and restarted.
Change the settings
The portal has no field for a function's scaling, CPU or memory. Use the CLI, with
functions/write:
stsh functions function update my-namespace hello -g my-resource-group \
--set 'scaling={"minReplicas":1,"maxReplicas":1,"cooldownPeriod":300}' \
--set 'resources={"cpu":"500m","memory":"512Mi"}'- Send all three scaling fields: one you leave out falls back to its default.
- The minimum must be 0 or more, the maximum at least 1 and not below the minimum.
- CPU and memory take Kubernetes quantities, such as
500mor1for CPU and512Mior1Gifor memory. A value that is not a valid quantity is refused.
The function is redeployed with the new settings.
Namespace defaults
The compute plan and the scaling you choose when you create a function namespace are not applied to its functions today. The plan is not saved at all; the scaling is saved as the namespace's default and shown on its Overview, but every function starts with the values in Settings a new function starts with, and changing the namespace's Configuration → Scaling does not change its functions.
The namespace's proxy
Each function namespace runs one replica of its proxy, fnproxy-<namespace>, which requests 0.1
CPU and 128 MiB of memory whether or not the namespace has functions. It never sleeps: it is what
receives calls, wakes sleeping functions, runs timer schedules and stops idle functions.