Skip to the essay
ShemolHibernate Container: A Deflated Container Mode for Fast Startup and High-density Deployment in Serverless Computing
容器 / 云原生 / 论文阅读 / Quark

Hibernate Container: A Deflated Container Mode for Fast Startup and High-density Deployment in Serverless Computing

🔗 https://github.com/QuarkContainer/Quark/blob/main/doc/Hibernate.pdf

English title: Hibernate Container: A Deflated Container Mode for Fast Startup and High-density Deployment in Serverless Computing

0 Abstract

Serverless computing is a popular cloud paradigm that needs low response latency to handle on-demand user requests. Two well-known techniques reduce that latency: keeping a fully initialized container alive (hot container1), or cutting the startup (cold start) delay of a new container. This paper proposes a third container-start mode: the hibernate container2. It starts faster than the cold-start mode and uses less memory than the hot-container mode. A hibernate container is essentially a “deflated” hot container. Application memory is swapped to disk, freed memory is reclaimed, and file-backed mmap memory is dropped. The deflated memory is inflated again when serving a user request. Because the application is already fully initialized, response latency is lower than cold start; because application memory is deflated, memory use is lower than a hot container. Further, when a hibernate container is “woken” to handle a request, the woken container has latency similar to a hot container, but uses less memory, since not all deflated memory needs to be inflated. We implemented hibernation as part of the open-source Quark secure container runtime. Our tests show hibernate-container memory is about 7% to 25% of a hot container. Altogether this means higher deployment density, lower latency, and a clear gain in overall system performance.

1 A hot container is a fully initialized container created as part of a hot-start path
2 Hibernate container refers to the hibernate-container mode

1 Introduction

Serverless computing, including Function as a Service (FaaS) and serverless containers, is becoming an increasingly popular cloud paradigm. Serverless environments typically run multi-tenant workloads in a shared setting to handle on-demand user requests. Major cloud providers support it: AWS Lambda [1]/Fargate [2], Google Function [3]/Cloud Run [4], and Azure Function[5]/Container Instance [6].

To host multi-tenant application workloads in a shared environment, major clouds usually use VM-based secure container runtimes rather than process-based ones such as runC/LXC. VM-based secure container runtimes give isolation like a traditional VM. For example, AWS uses Firecracker[7], GCP uses gVisor[8], and Alibaba Cloud and Huawei Cloud use Kata containers[9] for serverless. Those VM-based secure runtimes consume more memory and have higher response latency than process-based container runtimes.

User-request response latency is critical in serverless. It has three parts: container-runtime start, application init, and request handling. Runtime start is typically around 100 ms; application init ranges from 10 ms to 10 s; request handling is often short, a few milliseconds to 10 s. Compared with request handling, runtime start and application init contribute a lot. Cutting those two is one of the key design challenges. There are usually two optimization approaches:

  • Hot-start optimization: a common trick is to keep the execution runtime alive — a hot container — for a short time, so a later call of the same request can reuse it. Hot start does cut cold-start cost, but keeping containers alive burns a lot of compute, especially memory, which raises system resource demand. There is ongoing work on hot-start efficiency, e.g. cutting runtime overhead[10][8] and better keep-alive scheduling for hot containers[11].
  • Cold-start latency optimization: of course we cannot keep every serverless container alive under resource limits. Another line of work is reducing cold-start latency: faster runtime start[10][8] and faster application init[12][13][14].

This paper proposes a third start mechanism based on classic memory swapping, which fits serverless start-time quite well. Two related considerations:

  • Fast swap storage: with commercially available high-performance secondary storage (SSD, NVM) in public clouds[15], swap performance has improved a lot.
  • Lightweight serverless workloads: fast-start needs lightweight, low-footprint workloads. In AWS [16], 47% of functions use the default minimum 128 MB. Overall only 14% of AWS Lambda functions are allocated more than 512 MB. In Azure[17], 90% of applications never consume more than 400 MB, and 50% of serverless application workloads are allocated at most 170 MB. Small footprints mean relatively low swap cost.

So there is a real opening to use swapping for low-latency start plus low memory for keep-alive containers. With swapping as the key lever, this paper proposes and implements the hibernate container: a deflated keep-alive hot container. It uses these optimizations for low-latency start and low memory:

  • Memory: a hibernate container uses far less memory than a hot container because it swaps application memory to disk, reclaims free application memory back to the host kernel, and finally drops file-backed mmap memory back to the host OS.
  • CPU: it uses no CPU cycles, because the user application is fully paused.

User-request latency of a hibernate container is much lower than cold start, mainly because the application is fully initialized and the runtime is still using these keep-alive resources:

  • Host OS objects: the hibernate container keeps host OS objects alive — runtime process, cgroups, container network, container filesystem, processes. Those objects use little memory, but keeping them avoids a large re-init cost.
  • Blocked runtime threads: the runtime’s host threads block waiting for a user request. They use no CPU, but the system can respond immediately, like a hot container.

Importantly, a woken container’s latency on later requests is almost like a hot container, while using less memory, because it does not need all inflated memory to handle the request.

Overall this yields higher deployment density and better system performance.

Main contributions:

  • We propose and implement hibernate-container mode as part of the open-source Quark container runtime [18]. It uses less memory than a hot container and starts faster than cold start. A woken container derived from it also uses less memory than a hot container, with almost similar request latency.
  • We find the main swap-in delay is random SSD reads. Inspired by REAP[16] (a record-and-prefetch mechanism), we implement batched memory-prefetch swap-in as part of hibernate inflation. We compare page-fault swap-in and REAP swap-in across benchmarks.
  • We implement a new reclaim-oriented memory manager that efficiently returns free pages to the host kernel, avoiding the need for complex ballooning.
Ballooning pictures a balloon inside guest-occupied memory. Memory in the balloon is usable by the host (but not by the guest). When the host is tight on free memory, it can ask the guest to reclaim some of the memory allocated to it; the guest frees idle memory, and if that is not enough it may reclaim in-use memory and even swap some of it to the guest swap partition, inflating the balloon so the host can reuse that memory for other processes (or other guests). Conversely, when the guest is short on memory, the balloon can be deflated, releasing balloon memory back so the guest can use more.

2 Background and motivation

A hibernate container is a deflated hot container implemented in the Quark runtime. Deflation reclaims free application memory and swaps user memory to secondary storage. This section covers Quark’s secure-runtime design, existing guest-OS free-memory reclaim and swap, then our motivation and opportunities in today’s serverless stacks.

2.1 Secure containers and the Quark runtime

As noted, we implemented hibernate mode as part of open-source Quark [18]. Here we briefly survey state-of-the-art secure container runtimes, then go into Quark’s architecture.

Serverless hosts multi-tenant workloads in a shared environment. Traditional runtimes such as RunC/LXC are a poor fit, mainly because they cannot provide multi-tenant-level isolation. Major clouds instead use VM-level secure container runtimes: Kata[9]/Firecracker[19] and gVisor[8].

Figure
Figure

Kata and Firecracker use Linux-kernel VMs for isolation. Because both use a general-purpose Linux kernel, start latency and resource overhead are fairly high in serverless.

Quark[18] and gVisor[8] are two other notable secure runtimes designed for serverless. They consist of a user-space OS kernel plus a lightweight VMM. They aim to provide a Linux-compatible syscall interface and CRI/OCI, so existing Linux images can run unchanged. Both are highly optimized for serverless, so start latency and resource overhead are lower than Kata/Firecracker. Unlike Kata/Firecracker, which are explicitly based on the Linux kernel, Quark and gVisor’s Linux compatibility is not as good.

Figure 2 shows Quark’s architecture. It looks like a traditional Linux VM: it runs on a Linux host kernel with a KVM hypervisor. The Quark runtime process runs inside a standard Linux container, isolated with cgroups and network/filesystem namespaces. Quark includes a new user-space OS kernel (QKernel) and VMM (QVisor), both heavily optimized for serverless. Quark virtualizes a syscall interface that emulates Linux syscalls, and implements most Linux kernel functions: memory management, process management, I/O, and so on.

Quark is built for serverless and folds in serverless-specific features such as hibernate-container mode.

2.2 Reclaiming memory freed by the guest OS

A key value of hibernate mode is returning memory freed by the user application to the host OS. Reclaiming freed memory is not simple for a general-purpose guest OS like Linux. When a guest application frees memory to the guest kernel, ideally that kernel should return it to the host Linux kernel so it can be given to other host processes. Unfortunately the Linux guest keeps freed memory in its own pool instead of returning it. Linux is optimized for bare metal, where reclaim to a host is not required. In short: memory freed inside the guest is not released and reclaimed by the host in a traditional virtualization setup.

Two approaches address guest-Linux free-memory reclaim:

  • Ballooning[20]: relies on a special balloon driver in the guest, cooperating with the VMM to resize VM memory. Ballooning lets the hypervisor pull unused memory from some guests and share it with others.
  • Memory plug-in[21]: relies on kernel support for hot-add/hot-remove of physical memory. Hot-remove makes a region unavailable to users and needs page migration of in-use pages to another region, which costs performance. Memory plug-in has been used in VM memory collection [22].

VM-based runtimes such as Kata/Firecracker are also based on a Linux guest, so they have the same reclaim problem. To our knowledge neither uses ballooning or memory plug-in — those methods are too complex to adopt in serverless.

As part of hibernate-container work, we implemented a dedicated memory manager in Quark to raise reclaim efficiency in serverless.

2.3 Swapping guest application memory

Another key value of hibernate mode is swapping out user-application memory.

Swap temporarily writes inactive pages to secondary storage and marks the page-table entry not present. When the system needs that page, the VM subsystem raises a page fault to swap it in.

In a common virtualized setup, host swap is inefficient because swap is uncooperative. VSWAPPER[23] studies that inefficiency: silent swap writes, stale swap reads, false swap reads, and so on. VSWAPPER implements a guest-agnostic swapper that addresses those issues. It works reasonably for common virtualization, but is not particularly aimed at serverless.

For serverless we have several chances at better swap performance, because we want to swap out an idle container’s entire memory:

  • Batched swap-out of application pages: normally the kernel picks inactive pages. In serverless we swap the whole user-application working set of an idle container, which saves costs in the memory-management path.
  • Contention-free swap-out of a paused application: we can pause idle user processes while swapping, avoiding the messy races of ordinary swap.
  • Batched sequential disk reads for swap-in: ordinary page swap-in is page-fault driven and hits swap storage with random reads. For HDD or SSD, batched sequential reads always beat random reads. REAP[14] shows a function touches the same stable working-set pages across invocations. After identifying that set, those pages can be prefetched with batched sequential reads. Versus page-fault swap-in, batched swap-in saves both random disk-load cost and page-fault handling plus guest/host switch cost.

Our motivation is to take all of those opportunities and build a more efficient swap path for serverless.

3 Design and implementation

The following subsections go into the architecture and design of hibernate-container as part of the Quark secure runtime.

3.1 Hibernate-container state machine

This section describes how a hibernate container answers incoming user requests.

Figure 3 shows the container state transitions for serving incoming requests.

Figure
Figure

On an incoming request, the serverless platform does 1️⃣ a cold start. That produces a new hot container, and the request is forwarded to it. When the hot container receives the request it 2️⃣ moves to running to handle it, then 3️⃣ returns to hot when done.

To cut response latency, the platform may keep the hot container alive for a short time. If more requests arrive in that window (back-to-back), the hot container can serve them at low latency. While idle, though, it still occupies its application memory. Under memory pressure the platform may evict the hot container to free memory for other function containers. After eviction, the next request pays a higher cold-start delay. In short: more hot containers in the system, better request latency.

Besides those traditional states, we propose three new ones:

Hibernate: a hibernate container is a deflated hot container with a smaller footprint than a hot container. Instead of fully evicting a hot container, the platform can “deflate” it to hibernate to free memory.

The platform starts deflation by sending SIGSTOP to the hot container, moving it 4️⃣ from hot to hibernate.

Hibernate-running: on a user request, a hibernate container may 7️⃣ transition to hibernate-running, like a running container, to handle the request.

Woken: a hibernate-running container 8️⃣ returns to woken after the request finishes. The woken container 6️⃣ goes to hibernate-running again on the next request. It may also 9️⃣ go back to hibernate on SIGSTOP. Request latency of a woken container is almost like a hot container, with less memory. When the platform predicts an upcoming request, it may also “wake” a hibernate container into woken by sending SIGCONT 5️⃣, to cut latency.

3.2 Deflation Process Overview

A hibernate container is essentially a compressed hot container. It is derived from a hot container in four steps:

  1. Pause hot-container user-application processes, and block runtime host OS threads waiting for a “wake” trigger.
  2. Reclaim freed application pages and return them to the host Linux kernel.
  3. Swap committed application pages out to local disk.
  4. Drop file-backed mmap memory with madvise() using MADV DONTNEED as the advice, returning it to the Linux kernel.

After step #1 the hibernate container uses no CPU. After steps #2, #3, #4, allocated application memory has been returned to the host kernel, so it uses far less memory than a hot container. Details of #2, #3, #4 are in 3.3, 3.4, 3.5.

A hibernate container can become a hot container again via memory inflation plus resuming user processes. Two wake triggers:

  • User request: when the platform gets a request, it can forward it straight to the hibernate container without waking it first, to cut system latency. The hibernate container does this by blocking a runtime thread waiting on the request, e.g. Posix socket sys accept or sys read. When a client connect or socket data is ready, the host kernel unblocks the thread and the rest of wake processing runs: swap-in, then resume the application.
  • Serverless control plane: the platform can explicitly wake the container when a request is expected. Because inflation is partly done before the request arrives, latency is lower than the request-triggered path.

3.3 Reclaim-oriented memory management

A hibernate container reclaims pages freed by the guest application and returns them to the host Linux kernel, like VM ballooning.

QKernel runs in a KVM VM; its guest physical memory is virtual memory of the host Linux OS. Guest physical pages (host virtual pages) are not committed by the host kernel until accessed. Quark can return committed pages to the host with sys madvise() and MADV DONTNEED[24]. After a successful madvise(), later accesses in that range still succeed, but they cause zero-fill on-demand pages for anonymous private mappings. Ideally, once free regions are identified in the guest kernel allocator, they can be reclaimed with madvise().

Unfortunately the original Quark runtime could not easily reclaim freed memory. It currently uses a binary buddy allocator [25]. That is a poor fit for reclaim: free blocks live in a free list, a linear linked list whose “next” pointer is stored in the free block itself. For a bare-metal kernel that works. For a guest kernel, if we madvise() a free page block, later access is zero-filled, so the “next” pointer is cleared and the free list is corrupted. So the existing buddy allocator is not suitable for page reclaim. We implemented a bitmap page allocator instead; design follows.

Original Quark memory allocation has two domains:

  1. Guest user-application address-space allocator: the application gets address space from the guest kernel via syscalls (sys brk, sys mmap). Those only allocate address space; pages are not committed until a page-fault handler.
  2. Quark runtime global heap: Quark is written in Rust. The Rust runtime supports a custom heap allocator; Quark uses a buddy-based one. QKernel kernel data structures such as kernel stacks come from this heap. Original Quark also allocated user-application pages from the global heap in the page-fault handler, which is reclaim-unfriendly. So we added a third allocator for hibernate: the bitmap page allocator.
Figure
Figure

The bitmap page allocator manages page allocation for guest user applications. It is used only in the page-fault handler for fixed-size 4 KB pages.

As in Figure 4, it uses 4 MB chunks for 4 KB pages. The chunk start is 4 MB-aligned; the first 4 KB page is reserved as a control page with three fields:

  • “Next” pointer: all 4 MB chunks with free pages are linked in a free list (linear linked list). The “next” pointer lives in the control page.
  • Free-page bitmap: a 4 MB chunk has 1024 4 KB pages. The first is the control page; the remaining 1023 can be allocated. The allocator keeps an L2 bitmap, a 16 × 64-bit integer array (1024 bits), one bit per page allocated or not. To speed free-page lookup it keeps another 64-bit integer as an L1 bitmap, indicating whether each L2 64-bit word is zero. So each allocation touches two 64-bit integers: L1 and one L2 word. To find a free page, first find the first non-zero bit in L1.

Suppose that is bit 4: the 4th L2 word has a free page. Then find the first non-zero bit in that L2 word. Free-page lookup is thus O(2).

  • Page reference-count array: kernel pages may be referenced by multiple page tables (e.g. on clone). The allocator also keeps refcounts on the control page, a 16-bit atomic integer array[26], to maximize memory efficiency and performance.

Page allocation and refcount inc/dec:

  1. Allocation: on a request, allocate a page from the first 4 MB chunk on the free list and update the control-page bitmap. If that chunk has no more free pages, remove it from the list. If the list is empty, allocate another 4 MB chunk from the global heap (global buddy allocator). Allocation takes a global lock to avoid races.
  2. Refcount inc/dec: on guest process clone/exit or copy-on-write, guest page refcounts go up and down. Storing them on the control page helps performance. Because chunks are 4 MB-aligned, any guest page finds its control page by clearing the low 22 bits of its address (first page of the 4 MB chunk). No lookup table. After finding the control page, inc/dec uses Rust atomics: fetch_add / fetch_sub [26], lock-free. When refcount hits zero, the page is freed and returned via the free-page bitmap. If a chunk’s free-page count was zero and a page becomes free, the chunk goes back on the free list. When free-page count hits the max 1023, the 4 MB chunk can be returned to the global heap.

When Quark starts hibernate processing, freed pages must go back to the host kernel. Because free data pages are indicated by the control-page bitmap, the pages themselves do not store “next” pointers as in a buddy free list. At hibernate time Quark returns free pages with madvise(). That makes hibernate much simpler than VM ballooning.

Also, when a hibernate container is woken and reclaimed pages are reallocated to the guest application, the host kernel commits them via host page faults. That is transparent to guest Quark, so Quark need not explicitly reallocate pages. That cuts wake latency and system complexity.

3.4 Hibernate-container memory swap

Figure
Figure

A hibernate container swaps guest application memory to secondary storage; swapped pages may be swapped in when the container is woken.

Two swap-in mechanisms:

  • Page-fault swap-in: like ordinary OS swap-in, accessing a swapped-out page faults, and the fault handler loads the page.
  • Batched REAP swap-in: a hibernate container can prefetch all pages recorded by the REAP process.

As in Figure 5, each container sandbox has a swap file for page-fault swap-in and a REAP file for batched REAP swap-in. The swap file is dedicated to one sandbox and not shared, to reduce potential security issues; files are deleted when the sandbox exits.

Page-fault swap-in and REAP swap-out differ as follows.

3.4.1 Page-fault swap-in and swap-out

This section describes page-fault swap-in and the matching swap-out.

The serverless platform can trigger hibernate of an idle hot container with SIGSTOP. Then the swap manager in Figure 5 does:

  1. Pause the guest application: the swap manager pauses all guest applications as a generic SIGSTOP handler. After that, user threads are blocked and will not touch any page, so the swap manager need not handle complex races, simplifying the path;
  2. Walk and modify guest page tables: walk all guest application page tables, find anonymous pages, and:
    1. Mark each anonymous PTE not present, so a later access faults;
    2. Set PTE flag bit #9, a custom bit meaning the fault is due to swap-out;
    3. Put the page’s guest physical address in a hash table for dedup when multiple page tables reference the same GPA.
  3. Write pages to the swap file: each Quark sandbox has a swap file. The swap manager enumerates the hash table, writes pages to the file, and stores file offsets back in the hash table.
  4. Return pages to host Linux: madvise() the swapped-out pages.

When the hibernate container is woken, it resumes the paused application. Accessing a swapped-out page faults to swap it in.

Fault handling:

  1. Confirm the fault is from a swapped-out page: the Quark page handler checks custom PTE flag bit #9. If set, it is a swapped-out page; swap-in starts at step #2.
  2. Load memory from the swap file: the faulting vCPU exits from guest to host and reads the page from the swap file.
  3. Update the PTE: clear bit #9 and mark the entry present so this page will not fault again.

Page-fault swap-in is expensive. Costs include:

  • Page-fault handling: the guest vCPU moves from guest user space to guest kernel space; all general-purpose registers are stored in main memory.
  • Guest/host mode switch: expensive. Not only GPRs but floating-point context must be stored. In our test environment this guest/host switch is about 15 µs.
  • Random 4K reads from SSD: hibernate was tested on SSD. SSD random 4 KB read throughput is far above HDD, but still below sequential batched reads. In our setup, 4K random read is about 100 MB/s; sequential batched read exceeds 1 GB/s.

We observed that page-fault swap-in only loads 30% to 90% of swapped-out pages to handle a user request. For example, in our Node.js Hello World hibernate test, about 10 MB was swapped out, and request handling swapped in only about 4 MB. Swapped-out pages include both init and request-handling pages; when inflating a hibernate container, init is already done, so those pages are not touched. From that observation, and inspired by REAP[14], hibernate implements batched REAP prefetch swap-in.

3.4.2 REAP swap-out and batched swap-in (record and prefetch)

The main idea of REAP is to record the guest-physical working set while handling a user request, then prefetch it in batch on the next wake. It is an optimization of page-fault swap-in.

Versus page-fault swap-out, REAP adds a record pass after the first hibernate and wake. After the container first enters hibernate, recording does:

  1. Sample user request: the platform sends a sample request to trigger hibernate-running;
  2. Record the physical working set: in hibernate-running, the request process’s guest-physical working set is loaded from the swap file via page-fault swap-in; untouched pages stay in the swap file; 3. REAP hibernate: after the sample request, the container returns to woken. The platform sends SIGSTOP to hibernate the woken container. REAP swap-out:
    1. Pause all user processes.
    2. Walk all page tables to get all active anonymous pages.
    3. Record guest physical page addresses in a hashed I/O vector, and save pages to the REAP swap file in Figure 5 with batched pwritev().
    4. Release guest physical pages with madvise().

REAP swap-out differs from page-fault swap-out:

  • It does not change PTEs, so it does not trigger page faults;
  • It writes a dedicated REAP swap file, so swap-in can be accelerated by sequential batched disk reads.

REAP swap-in is simpler than page-fault swap-in.

Steps:

  1. Prefetch all pages from the REAP swap file with batched sequential preadv(), using the hashed I/O vector created at REAP swap-out.
  2. Resume guest application processes. REAP swap beats page-fault swap-in because:
    1. No page-fault overhead: REAP swap-out does not change PTEs, so swap-in does not fault — no fault handling or guest/host switch.
    2. Batched file reads: REAP swap-in prefetches all pages with batched reads, higher throughput than random reads.

3.5 File-backed memory sharing and security

A hibernate container also drops file-backed mmap memory with madvise() to return it to host Linux. Quark may share file-backed mmap memory across containers with copy-on-write. When memory is shared, deflation need not drop it, because other containers may still use it. Sharing can cut start latency and system-wide footprint — attractive for serverless.

Unfortunately, in multi-tenant serverless, sharing file-backed mmap across tenants creates security risk, e.g. cache side channels [27]. So memory sharing is not recommended in multi-tenant production [7].

For apps in a secure container, two main kinds of file-backed memory might be shared:

  • User language-runtime binaries: apps are built on runtimes such as Node.js and Python. Those files are mapped into user space and directly accessible. Sharing them across tenants is risky.
  • Secure-container runtime binaries: the runtime’s executable and libraries, e.g. Kata’s Linux guest kernel binary. These are not mapped into user space and the app cannot access them directly, so risk is lower than language-runtime binaries. RunD [10] shares the Linux guest kernel in production to speed cold start and cut footprint.

Hibernate enables sharing of Quark runtime binaries, but disables sharing of language-runtime binaries.

Runtime-binary sharing cuts request latency a lot. For example, we evaluated Node.js runtime binary sharing: with Node.js binary memory sharing enabled, hibernate Node.js hello-world request latency dropped from 25 ms to 11 ms. There are mitigations [28] [29] for language-runtime sharing risk; some are used in production. Cloudflare Worker [29], based on V8 isolates, accepts some multi-tenant isolation risk. Once those issues are resolved, hibernate may enable more file-backed sharing with those mitigations for better performance.

3.6 Hibernate-container implementation

The hibernate codebase is Rust, running as part of the Quark secure runtime. Quark is a virtual user-space OS with more than 200k lines of Rust. Hibernate changes key paths: virtual memory, I/O, and the VMM. We implemented the swap manager and bitmap allocator from scratch: 780 LOC for the swap manager, 484 for the bitmap allocator. The reclaim manager was built by changing Quark’s file-backed memory and page-table modules, about 500 lines. About 300 more lines changed in signal handling and I/O to trigger hibernate.

4 Evaluation

We show hibernate performance experimentally. Machine: 1×12-core Intel(R) Core(TM) i7-8700K @ 3.70 GHz, 64 GB RAM, PM981 NVMe Samsung 512 GB SSD, Ubuntu 20.04.4 Linux, kernel 5.15.0-46-generic.

Two sets of microbenchmarks measure the effect of hibernate on request latency and memory footprint.

  • Python benchmarks: we pick a set from Function Bench[30] covering different process types, memory use, and latency.
    • Floating-point: small memory and processing latency;
    • Video processing: applies OpenCV grayscale to video input. Footprint over 200 MB, processing latency over 1000 ms.
    • Image processing: image transforms with Python Pillow. We use two image sizes to see data size vs memory and latency of the same program.
  • Language-runtime hello-world: Python, Node.js, Golang, Java hello-world, to see different runtimes.

4.1 User-request response latency

This experiment shows hibernate request latency is lower than cold start, and a woken container is similar to hot. We also compare page-fault vs REAP swap-in latency.

Because a hibernate container is already started and keep-alive, we measure end-to-end application request/response latency, not application start latency. Except cold start, we run microbenchmarks as HTTP services, trigger processing with an HTTP request, and measure HTTP response latency.

Figure
Figure

We collected latency for cold start, hot, hibernate with page-fault/REAP swap-in, and woken:

  1. Cold start: process latency of container start plus request handling, without an HTTP trigger;
  2. Hot: request latency after the container is fully initialized;
  3. Hibernate: first-request latency after hibernate. We also collect page-fault and REAP swap-in latency.
  4. Woken: request latency of a woken container.

Figure 6 shows results. We conclude:

  1. Hibernate request-handling latency is lower than cold start: e.g. Hibernate REAP request latency is 3% (Python/Golang Hello-world) to 67% (image processing, 2.6 MB file) of cold-start process latency. Hibernate REAP saves 296 ms (Golang Hello-world) to 2407 ms (video processing) of cold-start process latency.
  2. Woken request-handling latency is almost like hot.
  3. Hibernate with page-fault swap has higher request latency than REAP: REAP beats page-fault swap on most benchmarks. The only exception is image processing at 2.6 MB, and the difference is negligible.

Results show hibernate response latency is below cold start, and woken handling latency matches hot. Hibernate also benefits language runtimes: Python, Node.js, Golang, and Java.

4.2 Container memory consumption

This experiment shows a hibernate container and its derived woken container use less memory than hot. We collect proportional set size (PSS) with Linux pmap for:

  • Hot: the container has handled a few user requests.
  • Hibernate: transitioned from hot to hibernate.
  • Woken: a hibernate container woken by a user request.

As in §3.4, hibernate shares Quark runtime binaries, so PSS is smaller when more instances run. We collected PSS for 10 running benchmark instances.

Figure
Figure

From Figure 7:

  1. Hibernate memory is far below hot: e.g. about 7% (video processing) to 25% (Golang Hello-world) of hot. Total savings from 12 MB (of 16 MB, Golang Hello-world) to 252 MB (of 281 MB, image processing, 2.6 MB file).
  2. Woken memory is below hot: e.g. 28% (Node.js Hello-world) to 90% (image processing, 2.6 MB) of hot. Total savings from 7 MB (of 16 MB Golang Hello-world) to 151 MB (of 226 MB video processing).

Results show hibernate and woken use less memory than hot, across Python, Node.js, Golang, and Java.

From latency and memory tests:

  1. Co-deploying hibernate and woken containers can achieve higher density than hot containers.
  2. A woken container uses less memory with request latency similar to hot. When possible, turning a hot container into a woken keep-alive via hibernate is beneficial.

5 Related work

Several related lines of work exist, especially cold-start optimization. Cold start usually has two activities: secure-runtime start and user-application start.

The next subsections describe ongoing work on both.

5.1 VM-based secure container runtime optimization

VM-based secure-runtime start includes Linux container environment setup (cgroups, container network, container filesystem) and VM OS kernel boot. RunD [10] speeds environment setup by pre-creating cgroups and optimizing container rootfs mapping. RunD uses Kata templates to cut per-microVM memory overhead and start latency.

Firecracker[19] introduces a lightweight VMM instead of QEMU or Cloud Hypervisor. It is aimed at serverless: few devices, virtio net and a single block device type. That helps secure-container memory use and sandbox start latency.

Quark[18] and gVisor[8] introduce a new user-space OS kernel and a new VMM aimed at serverless, cutting footprint and start latency.

5.2 Application-start optimization

The key idea in those cold-start optimizations is starting from a state “closer” to a hot container. [13][31] use checkpoint/restore from gVisor or the JVM. [32] reuses a hot container about to be reclaimed to host another image.

Based on checkpoint/restore (C/R), Catalyzer [13] implements init-less start. There are more optimizations on C/R, e.g. REAP uses batched prefetch to speed VMM image load. Sock [12] extends Zygote by forking a helper container with pre-imported packages to start a new container, saving import cost. Catalyzer introduces sandbox fork (sfork): fork from an existing hot container with full application state, sharing memory between parent and child.

Conclusion

Low-latency container start is critical to serverless user experience. Besides cold start and hot start in today’s serverless, this paper proposes a third mode: the hibernate container. It is essentially a deflated hot container.

Our experiments show hibernate memory is far below hot, and request latency is below cold start. When woken, the woken container also uses less memory than hot, with the same request latency. Altogether: higher deployment density, lower request latency, and a clear improvement in overall system performance.