JEP 545: Faster Startup and Warmup with ZGC
Summary
Improve application startup and warmup by enhancing the Z Garbage Collector to acquire and prepare physical memory more swiftly and efficiently in response to application needs.
Goals
-
Reduce startup time, i.e., the time to an application's first response.
-
Reduce warmup time, i.e., the time needed to reach peak performance.
Non-Goals
- It is not a goal to use the same memory management techniques on all operating systems. For example, we may use large pages by default on Linux but not on Windows.
Motivation
The Z Garbage Collector (ZGC), introduced in JDK 15, is designed for low latency and high scalability. ZGC does most of its work while application threads are running, pausing those threads only briefly. ZGC's pause times are independent of heap size: Applications can use heap sizes from a few hundred megabytes all the way up to multiple terabytes, always achieving low pause times.
The initial size of the heap, which can be set via the -Xms option, has a strong effect on both startup and warmup times. That is because, with ZGC, acquiring raw memory from the operating system and preparing it for use as heap memory is a costly operation. A small initial heap size makes for rapid startup but slow warmup, since ZGC must incrementally acquire and prepare more and more memory, on demand, as the application creates more and more objects. A large initial heap size, by contrast, makes for slow startup but rapid warmup, since ZGC acquires and prepares more memory up front, before the application starts.
In practice, the difference can be dramatic: On typical current hardware, when using ZGC, the HotSpot JVM starts up in about a tenth of a second with a 2MB initial heap but requires over ten seconds with a 32GB initial heap.
(You can further improve warmup time by specifying the -XX:+AlwaysPreTouch option, though this will further slow startup time.)
Ideally, you should not have to choose between rapid startup and rapid warmup. ZGC should be capable of expanding the heap rapidly from megabytes to terabytes while preserving low latency during startup, warmup, and beyond.
Description
We propose to enhance ZGC so that the HotSpot JVM
-
Starts up more quickly, by requesting just 2MB for the heap by default;
-
Warms up more quickly, by expanding the heap aggressively; and
-
Manages memory more efficiently, in units of 2MB rather than 4KB.
Together with adaptive heap sizing, these changes will make it neither necessary nor desirable to set the initial heap size (-Xms), set the maximum heap size (-Xmx), or specify the -XX:+AlwaysPreTouch option. When running your application, you need only choose ZGC and enable adaptive heap sizing:
$ java -XX:+UseZGC -XX:+ZAdaptiveHeapSizing ...The benefit of these improvements is significant: In the case study below, which examines a prototypical server application, these changes reduce startup time by almost 40% and warmup time by almost 90%.
Before describing the enhancements in detail, we first review the fundamentals of modern memory management.
A primer on modern memory management
The Java Virtual Machine presents the illusion of infinite heap memory. The role of a garbage collector such as ZGC is to implement that illusion with the finite physical memory of the computer on which the JVM is running, using the facilities of the hardware and the operating system.
Virtual and physical memory
The HotSpot JVM runs in an operating-system (OS) process with a 64-bit virtual address space. It cannot access physical memory directly; rather, the CPU translates virtual addresses to physical addresses at each memory access.
The CPU's memory-management unit divides the virtual and physical address spaces into pages, typically 4KB in size. The OS assists the CPU in the translation of virtual addresses to physical addresses by maintaining a page table that maps virtual pages to physical pages. This mapping can be arbitrary; consecutive virtual pages are not necessarily mapped to consecutive physical pages. ZGC uses OS APIs to manipulate the page table and thereby control how the heap, in virtual pages of the JVM process, is mapped to physical pages.
Reserving virtual memory
At startup, ZGC reserves a contiguous range in the virtual address space that is 16 times the maximum heap size (-Xmx). If, e.g., the maximum heap size is 1GB, ZGC reserves 16GB of virtual memory, which works out to 4M virtual pages. Reserving virtual pages is mere bookkeeping by the OS, so it is fast. This operation does not record a mapping from virtual to physical pages in the page table. (A vast virtual address space allows ZGC to better manage fragmentation induced by large objects, discussed below.)
Committing physical memory
At startup, no physical memory is assigned to any of the vast number of reserved virtual pages. To acquire initial physical memory for the heap, ZGC asks the OS to commit a number of physical pages to the JVM process. That number corresponds to the initial heap size (-Xms). If, e.g., the initial heap size is 2MB, ZGC creates a heap comprising 512 virtual pages by asking the OS to commit 512 physical pages.
In contrast to reserving virtual pages, committing physical pages is no mere bookkeeping. The OS must first find unused physical pages, which may require swapping virtual pages of other processes out to disk or, in the worst case, killing other processes. Once the OS finds some physical pages, it zeroes them in order to avoid data leaks. Only then can the OS map those pages to virtual pages in the heap.
Expanding the heap
As a Java application runs, it creates objects in the heap. At some point, the application may create an object for which no space can be found. To avoid throwing an OutOfMemoryError, ZGC expands the heap by asking the OS to commit more physical pages to the JVM process. If, e.g., ZGC expands the heap from 2MB to 16MB, it needs 3,584 new virtual pages (from the 4M reserved earlier) to map to 3,584 new physical pages.
As with initial heap creation, heap expansion can be slow because committing physical pages is no mere bookkeeping.
Contracting the heap
ZGC groups the virtual pages of the heap into regions of between 2MB and 32MB. As ZGC runs, it detects when a region no longer contains live objects. Such a region, which now contains only garbage, is said to be evacuated. The virtual pages of an evacuated region are no longer needed; this could be as many as 8,192 virtual pages in a 32MB region.
The virtual pages of evacuated regions may be used to satisfy subsequent object allocations or, alternatively, be released back to the OS, effectively contracting the heap. Contracting the heap will increase the frequency of garbage collection, so ZGC heuristically chooses to contract the heap when both heap usage and garbage collection frequency are low. To contract the heap, ZGC asks the OS to unmap the physical pages that are presently assigned to evacuated virtual pages and then uncommit the physical pages from the JVM process.
Harvesting memory
ZGC can repurpose the physical pages committed to evacuated regions in order to speed up object creation. For example, the expression new int[1024*1024] creates an array that requires 4MB of memory, for which ZGC must find contiguous space in the heap. Even if the heap has that much free space in total, the space could be scattered across discontiguous virtual pages. In this situation, most garbage collectors compact the heap by moving objects from one region to another, laboriously trying to construct a region with at least 1,024 contiguous empty virtual pages of 4KB each. However, this generally requires a garbage collection cycle — walking the heap, marking live objects, and relocating them — which is expensive.
ZGC instead uses an OS facility, remapping, to construct a region with sufficient contiguous space. Remapping modifies the page table to map some virtual pages to physical pages that previously backed other virtual pages. An evacuated region of the heap holds only garbage, so the virtual pages comprising the region are no longer needed; the physical pages to which those virtual pages are mapped can safely be used for other virtual pages. Accordingly, ZGC harvests those physical pages by mapping a contiguous range of virtual pages, in the vast virtual address space that it reserved at startup, to those physical pages.
Harvesting enables ZGC to find space for a large object several orders of magnitude faster than doing a garbage collection cycle to compact the heap. Harvesting does not change the JVM's memory usage, because no new physical pages are committed — previously committed physical pages are merely repurposed.
Faster startup with a small initial heap
Traditionally, ZGC committed 1/64 of physical memory for the initial heap at startup. This policy resulted in, e.g., a 16MB initial heap on a machine with 1GB of physical memory, and a 2GB initial heap on a machine with 128GB of physical memory.
Generally, applications do not detect how much memory is available at startup and adapt their allocation to it. Instead, deployers make empirical observations of the application's performance with various initial heap sizes and then settle on a value for the -Xms option in the application's startup script, perhaps choosing to trade off faster startup for slower warmup, or vice versa, as described earlier.
Assessing the impact of initial heap size on startup is straightforward: Deployers can easily measure the time from starting the JVM to the application's first response. Assessing the impact of initial heap size on warmup, however, is harder: Deployers must measure the time required not just for the first response but for potentially many responses, until the response time becomes stable. This is rarely done in practice, so it would be best to avoid the need for such assessments in the first place.
We propose that ZGC commit just 2MB for the initial heap, by default. Since committing memory is more than mere bookkeeping for the OS, it can be slow, so committing as little as possible — enough for only one 2MB region — reduces startup time.
Improved warmup with aggressive heap expansion
Applications will start quickly with a small initial heap, but usually that will be insufficient. The JVM will trigger more and more frequent garbage collection cycles as the application creates more and more objects. Collecting only a small heap means these cycles will be short, and ZGC uses dedicated worker threads to collect garbage without blocking application threads. Eventually, however, ZGC's rate of garbage collection will not be able to keep up with the application's rate of object creation. In that case, either ZGC must commit more physical memory, or the JVM must pause object creation until the end of the next collection cycle; either way, the application will pause.
We propose that ZGC expand the heap more aggressively than it does today. When it first detects a high frequency of garbage collection cycles, it will proactively expand the heap. Over time, ZGC will estimate the application's normal allocation rate and use that to calculate when to proactively expand the heap. ZGC will rapidly expand the heap by successively larger amounts; this will ensure a logarithmic bound on the number of expansions, and thus the number of collection cycles, needed to achieve a stable heap size. Finally, ZGC will expand the heap concurrently, using its own worker threads rather than application threads.
Aggressive heap expansion will allow ZGC to reduce the frequency of collections, in turn reducing the CPU overhead of garbage collection and improving the throughput of the application. The time spent waiting for the OS to commit physical pages will be incurred by ZGC worker threads rather than by application threads, resulting in more consistent execution and fewer pauses. To ensure that application threads are never starved, ZGC will never use more than 25% of the available CPU threads for its worker threads.
Faster warmup with concurrent pre-touching
When a process asks the OS to map a range of virtual pages to a range of physical pages, the OS does not immediately update the page table to reflect that mapping. Doing so could be a waste of effort, since some processes never use all their virtual pages. The OS therefore updates the page table incrementally, on demand: The first time that a newly mapped virtual page is read or written, the CPU raises a page fault, which causes the OS to update the page table to map that page to the corresponding physical page.
In a JVM process, page faults frequently occur during warmup, as the application creates many objects and the JVM writes to many virtual pages for the first time. To improve warmup in JDK 25, we suggested using the -XX:+AlwaysPreTouch option, which causes ZGC to pre-touch the initial heap, i.e., write to each of its virtual pages immediately after mapping them to newly committed physical pages, thereby causing the OS to update the page table immediately.
The cost of this approach is that JVM startup time increases, since all of the page faults for the initial heap now occur during startup. If you set the initial heap size (-Xms) to a small value, you might not notice an impact, but if you set the initial heap size to the same as the maximum heap size (-Xmx), you will likely see much slower startup. Here are some comparative wall-clock times for JVM startup with -XX:+AlwaysPreTouch and various initial and maximum heap sizes:
-Xms2M -Xmx2M 0.034s
-Xms20M -Xmx20M 0.041s
-Xms200M -Xmx200M 0.110s
-Xms2G -Xmx2G 0.634s
-Xms20G -Xmx20G 5.672sWe propose that ZGC pre-touch the heap not at startup but, rather, when it expands the heap. Whenever a ZGC worker thread expands the heap by committing more physical pages, it will write to the newly mapped virtual pages, forcing the OS to update the page table for those pages immediately rather than update them incrementally as they are first used.
Faster memory access via large pages
Modern applications routinely process gigabytes of data, so organizing virtual memory in units of 4KB pages is not efficient. Linux and Windows support large pages, sometimes known as huge pages, which are typically 2MB in size. An 8MB array, e.g., can be stored in four large pages instead of 2,048 small pages.
Having fewer but larger pages makes for faster access. Because the virtual address space is huge, the page table is not a simple lookup table but, rather, a hierarchy of tables. When the JVM dereferences a virtual address, the CPU looks through successive levels of the hierarchy to determine a physical address. Using large pages allows the CPU to skip one level of lookup. Additionally, the CPU maintains a cache of recent virtual-to-physical mappings, and a large page occupies only one cache entry versus 512 entries for the corresponding small pages; this makes a cache hit more likely for a larger range of addresses.
Historically, large pages were disabled by default in the Linux kernel. Enabling large pages required the superuser to configure the kernel, but this was not practical in many deployment environments.
As of Linux 6.1, released in 2022, any process can now ask the OS to upgrade a specified range of small virtual pages into large pages; no kernel configuration is required. ZGC will use this API to upgrade small pages to large pages during heap expansion so that, eventually, the entire heap is backed by large pages. Since upgrading can be a costly operation, potentially requiring physical memory to be defragmented, it will be performed by ZGC worker threads so as to minimize impact on the application.
Case study in faster startup and warmup
We explore the startup and warmup of the Spring PetClinic application on a machine with 32GB of physical memory.
To quantify warmup, we measure the response time of the application in the period immediately following JVM startup. Specifically, we measure P99 response times, i.e., 99th-percentile response times, which characterize tail latency. We assume that an acceptable response time for this application is 20ms.
We use a custom load-generation tool to load the application. The tool initially generates a single request to simulate the health check performed by a load balancer. It then generates 1,000 requests per second, increasing linearly to 30,000 requests per second over a ten-second ramp-up period. The ramp-up period simulates a load balancer configured under the assumption that the application does not achieve peak performance until the JVM has JIT-compiled all hot paths to machine code. If the application were to receive 30,000 requests per second immediately, then P99 response times would be measured in seconds rather than milliseconds. By gradually increasing the request rate, we can achieve acceptable latency throughout. After the ramp-up period, the load generator continues to send 30,000 requests per second for two minutes.
We consider the application to be warmed up once the ramp-up period has finished and P99 response times consistently stay below 20ms.
Scenario 1: JDK 25 with a small initial heap and no pre-touching
To minimize startup time, we configure the JDK 25 JVM with a small initial heap (2MB) and a large maximum heap size (20GB):
$ java -XX:+UseZGC -Xms2M -Xmx20G -jar petclinic.jarThe time from JVM startup to the first response is 5.86s. Unfortunately, the P99 response times repeatedly spike above 20ms for tens of seconds since ZGC must expand the heap from 2MB to 20GB as the application demands more memory. That involves incrementally committing physical pages and triggering a page fault on the first access to each of the corresponding virtual pages, leading to poor latency. As the seconds tick by, allocations are increasingly satisfied in previously committed memory recycled by ZGC, rather than by newly committed memory, and P99 response times above 20ms become less frequent. By 40 seconds into the run, P99 spikes are rare, and by about 80 seconds, the application has warmed up.
Scenario 2: JDK 25 with a large, pre-touched initial heap
To improve warmup, we configure the JDK 25 JVM with a large initial heap (20GB) and ensure that it is pre-touched:
$ java -XX:+UseZGC -Xms20G -Xmx20G -XX:+AlwaysPreTouch -jar petclinic.jarStartup is worse due to the time needed to commit physical pages for the large initial heap and pre-touch the corresponding virtual pages — 9.45s rather than 5.86s — but warmup is visibly improved due to having plenty of committed and touched physical pages. The P99 response times are never above 20ms, so after the initial 10s ramp-up period, the application has warmed up.
Scenario 3: JDK NN
The enhancements to ZGC described above combine the startup benefit of a small initial heap, the warmup benefit of aggressive heap expansion and concurrent pre-touching, and the general benefit of large pages:
$ java -XX:+UseZGC -XX:+ZAdaptiveHeapSizing -jar petclinic.jarStartup takes 5.84s, which is essentially the same as scenario 1 but 38% lower than scenario 2. As in scenario 2, the P99 response times are never above 20ms, so after the initial 10s ramp-up period, the application has warmed up.
The improvement in warmup time from 80s in scenario 1 to 10s in scenario 3, while maintaining low startup time, can be attributed to the combination of aggressive heap expansion, concurrent pre-touching, and the use of large pages. We can demonstrate this by enabling and disabling the features individually:
-
If we rerun scenario 3 with aggressive expansion and large pages but without concurrent pre-touching, then warmup takes 34s.
-
If we rerun scenario 3 with aggressive expansion and concurrent pre-touching but without large pages, then warmup takes 18s.
All three features, working together, reduce warmup time so that it occurs entirely within the 10s ramp-up period. This is significant: When the 99th-percentile response time remains low, more CPU resources are available for the JVM to JIT-compile the application's hot paths to machine code.
Future Work
-
The key to fast warmup is reducing heap usage by spending CPU resources on concurrent garbage collection while, at the same time, increasing heap capacity by spending CPU resources on concurrently committing memory. We will investigate how to improve warmup further by more explicitly balancing CPU resources between these two activities.
-
Ahead-of-time code compilation (JEP 544) from Project Leyden will remove the overhead of JIT-compiling the application to machine code during warmup. This will make more CPU resources available to ZGC for either concurrent garbage collection or concurrent memory commitment, or both.
Alternatives
Instead of committing 2MB for the initial heap, ZGC could commit a larger amount, which would reduce the need for aggressive expansion. However, if the user has said nothing about the initial heap size then it is reasonable not to assume it should be larger. Some applications create very few objects, and a larger initial heap would not be useful to them. The disadvantages of starting with an initial heap that is too small for an application are mitigated by committing and pre-touching memory for the heap while the application is running, and by expanding the heap exponentially.
Risks and Assumptions
-
A risk of the default initial heap size being 2MB rather than a function of the amount of physical memory is that applications may have been tuned to expect a larger initial heap size in their deployment environments. For example, an application tuned to run on a machine with 128GB of physical memory, where the initial heap defaults to 2GB in JDK 25, may experience slower warmup on JDK NN with a 2MB initial heap. The slower warmup may manifest as a prolonged period after startup during which P99 response times routinely exceed acceptable levels. We assume that deployers are familiar with the need to re-tune their GC configurations after significant enhancements to JVM functionality.
-
Many applications are started by scripts that configure the JVM with a small initial heap, far less than 2GB. These applications will likely warm up more quickly on JDK NN than JDK 25. Across the Java ecosystem, we assume that the enhancements to ZGC proposed here will be a net positive for both startup and warmup.