Chart with a steadily rising RSS line next to a flat GC heap line, surrounded by a grid of leaked 3400-byte memory blocks
Back to blog
.NET 10Memory LeakLinux

3.4 KB at a Time: How We Found a Native Memory Leak in .NET 10 on Linux

Sven HennessenDevelopment

TL;DR

  • Symptom: the RSS of each pod grows linearly by about 22-35 MiB per day, independent of load, while the managed GC heap stays small and flat.
  • Cause: on Linux, the .NET 10 runtime loses about 3.4 KB each time it rethrows a cached TypeInitializationException from native code.
  • Trigger: Npgsql 10 defaults to GssEncryptionMode=Prefer; without libgssapi_krb5.so.2 in the image, every new physical PostgreSQL connection causes one such rethrow.
  • Mitigation: set GssEncryptionMode=Disable in the connection string, set PGGSSENCMODE=disable, or add libgssapi_krb5.so.2 to the image (e.g. apk add krb5-libs). A runtime fix is in review (dotnet/runtime#135290).

Some of our ASP.NET Core services on Kubernetes used more memory each day while the managed heap stayed small. This post walks through how we traced the growth to a native memory leak in the .NET 10 runtime on Linux, why a Npgsql default triggers it, and three ways to stop it.

Some of our ASP.NET Core services on Kubernetes used more memory each day. The managed heap stayed small. Restarts made the problem go away for some weeks.

This post shows how we found the cause. The cause is a native memory leak in the .NET 10 runtime on Linux. A database driver setting triggers the leak, and either one configuration value or installing the missing native library stops it.

Why this is important for operations and cost

A slow leak does not cause an incident on the first day. It causes these problems:

  • The memory requests become too large. Teams set the memory request for the size of the pod after some weeks, not for the real working set. In each namespace, this uses quota that other workloads need.
  • Restarts that nobody plans. When a pod gets to its memory limit, Kubernetes stops it (OOMKill). Requests that are in progress fail.
  • The problem stays hidden. Each deployment restarts the pods. Teams that deploy frequently never see the problem. Teams that deploy less frequently get the OOMKills.
  • Alerts with no clear cause. The memory alerts start, but the usual .NET tools show a healthy managed heap. The team must then decide whether to investigate or to schedule restarts.

In our case, a pod with a 2 GiB limit would get to the limit after some weeks.

Step 1: detect the pattern

Our monitoring sent memory alerts for one service in production. We examined the RSS of all pods over 14 days:

ObservationValue
RSS of each podLinear growth, 22–35 MiB/day (most pods about 28–29 MiB/day)
Affected podsAll production pods of the service
Relation to loadNone. Idle pods grew at the same rate as busy pods.
Managed heap (Gen2, LOH)Small and stable
Thread countStable
Other services on the same platformSome grew at the same rate. Others stayed flat.

Two facts were important:

  1. The growth is linear and does not depend on load. A leak that the load drives grows faster on busy pods. This leak grew at the same rate on all pods. Thus a timer or a periodic background task probably causes it.
  2. Some services stay flat. These services use the same base image, the same platform, and the same chart. We used them as control services in each next step.

Step 2: prove that the managed heap is not the cause

The usual first step is a GC heap analysis. But first we had to know whether the managed heap could explain the growth.

Our monitoring showed only a small number of .NET runtime values for these pods. Thus we enabled the .NET runtime metrics in the service and sent them to Prometheus. The metrics include the GC heap size for each generation, the GC count, the thread pool, the exception count, and the process memory (process_working_set_bytes, process_private_memory_bytes).

Result:

  • The GC heap stayed small and stable.
  • The RSS continued to grow.
  • The difference between RSS and the GC heap grew linearly.

This difference is the important metric. If the GC heap is flat and the RSS grows, a managed memory profiler does not find the leak. The memory is in native allocations.

Tip: Put process_working_set_bytes and the GC heap size on the same chart. If the gap between them grows, stop the managed heap analysis and examine the native memory.

Step 3: find which memory regions grow

On Linux, /proc/<pid>/smaps shows each memory mapping of a process and its RSS. You can read it with kubectl exec. You do not need a debugger, and the pod continues to run.

We grouped the mappings into categories with a small awk script. A snapshot of one pod after 34.5 hours:

ConsumerRSSShare
glibc malloc thread arenas136.5 MiB24.8 %
glibc main arena ([heap])97.8 MiB17.8 %
glibc total234.3 MiB42.5 %
JIT code mappings (/memfd:doublemapper, W^X)113.6 MiB20.6 %
Other anonymous memory (managed heap, stacks)67.0 MiB12.2 %
Other libraries and assemblies98.0 MiB17.8 %

The managed heap was only about 14 MiB. The largest consumer was memory that native code allocated with malloc.

The control service changed our conclusion. A first theory was "glibc arena fragmentation". This is a known problem for .NET on Linux, and MALLOC_ARENA_MAX is the usual answer. But the flat control service had the same number of arenas (15). Its glibc share of RSS was also the same (42.6 %). Thus the arena count did not cause the growth. Fragmentation explains why glibc keeps the memory. It does not explain why the process continues to allocate more.

Lesson: MALLOC_ARENA_MAX and DOTNET_EnableWriteXorExecute=0 can make the RSS smaller. They do not fix a leak. Compare with a control service before you change these settings.

Thus we had to know what was in the growing malloc blocks.

Step 4: collect a dump from a running pod

We used dotnet-dump collect --type Heap in a dev pod after five days of uptime.

Some practical notes:

  • Use --type Heap, not Full. The heap dump took about 13 seconds. The liveness probe allowed about 30 seconds. A full dump takes longer and can cause a restart.
  • Copy the single-file dotnet-dump binary into the pod. Download it from https://aka.ms/dotnet-dump/linux-x64. Our runtime image did not contain it.
  • Set the extract directory with DOTNET_BUNDLE_EXTRACT_BASE_DIR to a writable path. If the image is in invariant globalization mode, also set DOTNET_SYSTEM_GLOBALIZATION_INVARIANT=1.
  • Do not trust kubectl exec ... cat for large files. The stream was truncated without an error. We compressed the dump, split it into parts of 20 MB, and copied each part in a loop until its SHA-256 checksum matched.

The result was a core file of 1.34 GB (122 MB compressed). The pod did not restart.

Step 5: examine the native heap

dotnet-dump analyze (SOS) shows the managed heap. For the glibc heap we used chap. chap reads a Linux core file and lists each malloc allocation. It also shows if an allocation is still referenced or is leaked.

Warning: our image used glibc 2.44. chap did not support this version completely and showed warnings about the arena layout. Thus we used the chap numbers only as a starting point, and we verified the important values with our own tools.

The result was very clear:

FindingValue
Allocations of exactly 3400 bytes (0xd48)about 44,600
Of these, not referenced by any memory (leaked)42,231
Total sizeabout 152 MB
Calculated growthabout 28 MiB/day

The calculated rate agreed with the RSS growth in our monitoring. Thus one type of allocation caused all the growth.

What is in a 3400-byte block?

chap could not identify the blocks. Thus we wrote a small Python script that reads the core file through its ELF PT_LOAD segments. The script calculated statistics for each 8-byte field over all 42,231 blocks.

Some fields had the same value in all blocks:

OffsetValueMeaning
0x300x10000FContextFlags (x64 CONTEXT_FULL)
0x380x33SegCs: the 64-bit user code segment
0x440x246EFlags: a typical flag value
0x1000x37Fx87 FPU control word (default value)
0xF8one addressRip: the instruction pointer

This is the layout of a Windows x64 CONTEXT structure, which is a snapshot of the CPU registers. The .NET runtime uses this layout on Linux too (in the PAL). With the AVX-512 fields, a CONTEXT and an EXCEPTION_RECORD fit into 3400 bytes.

The Rsp field (stack pointer) had 82 different values. Thus many different threads created the blocks. But Rip was the same in all blocks. Thus one single location in the code created all of them.

Which code created the blocks?

We found the load address of libcoreclr.so in the dump. We subtracted it from Rip and got an offset in the library. Then we disassembled the library at that offset. We used the same libcoreclr.so file as the pod and checked its SHA-256.

The address was the return address after a call to RtlCaptureContext. We compared the function with the .NET 10 source code and found it: DispatchManagedException(PAL_SEHException& ex, bool isHardwareException) in vm/exceptionhandling.cpp.

This function moves an exception from native runtime code into the managed exception handling. Thus each leaked block represents one exception that the runtime threw from native code into managed code.

Step 6: find which exception

Now we had to find the exception. We opened the same dump with dotnet-dump analyze and SOS.

  • The GC committed memory was 55.7 MB. This confirmed again that the managed heap was not the leak.
  • The heap contained a cached TypeInitializationException for the internal type Interop+NetSecurityNative. The inner exception came from the static constructor of GssInitializer. The GSSAPI native library was not in our container image, and thus its initialization failed.
  • The stack of the exception was:
InitHelpers.InitClassSlow
NetSecurityNative.ImportPrincipalName
SafeGssNameHandle.CreateTarget
UnixNegotiateAuthenticationPal.InitializeSecurityContext
NegotiateAuthentication.GetOutgoingBlob
Npgsql.Internal.NpgsqlConnector.GSSEncrypt

The TypeInitializationException behaves like this: if a static constructor fails, the runtime keeps the exception. Each later access to the type rethrows the same exception. The runtime does this rethrow from native code (InitClassSlow). Thus each access goes through DispatchManagedException.

Step 7: connect the two sides

Npgsql side. In Npgsql 10 the default GssEncryptionMode is Prefer. When SSL is on, Npgsql tries GSS encryption first on each new physical connection. It calls NegotiateAuthentication.GetOutgoingBlob, catches the TypeInitializationException, and continues with TLS. The connection works, and the application does not see an error. But each new physical connection causes one native rethrow.

Our service opens new physical connections regularly. Some data sources have no minimum pool size, and Npgsql closes idle connections.

Runtime side. In .NET 10 on Linux, the macro UNINSTALL_MANAGED_EXCEPTION_DISPATCHER_EX (in vm/exceptmacros.h) catches the native PAL_SEHException. It moves the exception into a local copy and calls DispatchManagedException. This call does not return. The managed exception handling continues in the managed catch block. Thus the destructor of the local copy does not run, and the destructor does not release the exception records (CONTEXT + EXCEPTION_RECORD). The runtime allocated these records with posix_memalign. Each rethrow loses one 3400-byte block.

Step 8: isolate the problem with a repro

A theory from a dump is not proof. Thus we wrote two small console applications and ran them on .NET 10.0.12 in a container image without the GSSAPI library. We used the Microsoft image mcr.microsoft.com/dotnet/runtime:10.0-noble-chiseled. This let us check whether the problem was specific to our own image.

Repro 1: runtime only, no database

The application calls NegotiateAuthentication.GetOutgoingBlob in a loop and catches the exception.

ImageCallsRSS growthFor each call
mcr.microsoft.com/dotnet/runtime:10.0-noble-chiseled40,000152 MBabout 3.9 KB
Control: managed throw/catch40,00014–15 MB, not linearn/a

Repro 2: end to end with Npgsql and PostgreSQL

The application opens Npgsql 10.0.3 connections to PostgreSQL 17 with SslMode=Require and Pooling=false. We compared Prefer and Disable.

ImageGssEncryptionModeConnectionsRSS growthTypeInitializationExceptions
Microsoft chiseledPrefer30,000103 MB30,000
Microsoft chiseledDisable30,0005 MB0

Results:

  • With Prefer, each connection causes exactly one exception. The RSS grows linearly by about 3.4 KB for each connection.
  • With Disable, no exception occurs. The RSS stays flat.
  • The leak occurs in the unmodified Microsoft image. Thus it is not specific to our own image.

The mitigation

There are three ways to remove the trigger. Choose one.

Option 1: disable GSS encryption in the connection string. This is what we did. We made GssEncryptionMode configurable in the service and set it to Disable through an environment variable in the Helm values:

Host=db;Database=app;SslMode=Require;GssEncryptionMode=Disable

Option 2: disable GSS encryption with an environment variable. If you cannot change code, Npgsql also reads the standard libpq variable:

PGGSSENCMODE=disable

Option 3: add the native GSSAPI library to the image. If libgssapi_krb5.so.2 is present, the static constructor of GssInitializer succeeds. Thus no TypeInitializationException is cached, and the native rethrow does not occur. On an Alpine-based image, add one line to the Dockerfile:

RUN apk add --no-cache krb5-libs

On a Debian- or Ubuntu-based image, the package is libgssapi-krb5-2 (apt-get install -y --no-install-recommends libgssapi-krb5-2). Chiseled images have no package manager. For these, you must copy the library in from a build stage, or use option 1 or 2.

Before you apply one of these options, examine these effects:

  • No functional change for us. GSS encryption never worked in our containers, because the GSSAPI library was not in the image.
  • TLS stays on. SslMode does not change.
  • Options 1 and 2 disable Kerberos for PostgreSQL. If you need Kerberos or GSS authentication, use option 3.
  • Option 3 makes the image larger and keeps the GSS attempt. Npgsql still tries GSS encryption on each new physical connection, but the attempt no longer goes through the leaking native rethrow.
  • The runtime bug stays. Other code that rethrows a cached TypeInitializationException from native code can also leak memory. These options remove only the trigger that we found.

Result

We deployed the mitigation to our development environment first. On the next morning, the RSS of the service was flat. Before the change, it grew by about 28 MiB per day.

How to check whether your services are affected

  1. Examine the RSS trend over some days. Look for linear growth that does not depend on load, while the GC heap stays flat.
  2. Count the exceptions. With .NET runtime metrics or dotnet-counters monitor System.Runtime, look for a constant rate of exceptions when the service is idle.
  3. Find the trigger. Run dotnet-trace or use a first-chance exception handler to log TypeInitializationException. Or collect a heap dump and run dumpheap -type TypeInitializationException in SOS.
  4. Check the image. Does your image use Npgsql 10 with SSL and no libgssapi_krb5.so.2? Then you probably have the trigger.

Upstream status

  • .NET runtime: we reported the leak with the repro in dotnet/runtime#135234. A fix is in review in dotnet/runtime#135290. The fix lets the exception dispatch take ownership of the exception records, and it adds a regression test. When we wrote this post, the fix targeted the main branch. A backport to .NET 10 was not announced.
  • Npgsql: npgsql/npgsql#6416 reports the same exception storm and memory growth in containers. Npgsql 10.0.2 catches the exception, but the exception still occurs on each new connection. npgsql/npgsql#6593 stops the GSS attempts after the first failure. This removes the trigger without a configuration change.

Lessons

  1. Measure the gap between RSS and the GC heap first. It tells you if the problem is managed or native memory.
  2. Always use a control service. Our first theory (glibc arenas) was wrong. The comparison with a flat service showed this.
  3. A leak that does not depend on load points to a timer or a periodic task. Find the period and match it to the code.
  4. Group native leaks by allocation size. One exact size that occurs tens of thousands of times is a fingerprint.
  5. The contents of a leaked block tell you who allocated it. Constant field values identify the structure. The instruction pointer identifies the code.
  6. An exception that the application catches is not free. It costs CPU. In this case, it also cost memory.
  7. Confirm the theory with a minimal repro, and include the vendor's own image. This makes the upstream report quick to accept.

Tools that we used

  • Prometheus and .NET runtime metrics
  • /proc/<pid>/smaps with awk
  • dotnet-dump (collect and analyze, SOS)
  • chap for the glibc heap
  • A small Python script that reads ELF core files
  • objdump for the disassembly of libcoreclr.so
  • The .NET runtime and Npgsql source code on GitHub
  • Podman for the repro containers

Sources

Need support?

The leak wasn't in our code but in a default nobody had consciously chosen: Npgsql's GSS setting, combined with a native library missing from the image. These are exactly the spots we look at when we review .NET services: connection strings and driver defaults, base images and Dockerfiles, resource limits and everything else that only surfaces after weeks in production. A focused review of your service shows you which defaults are quietly running in yours.

Get your service reviewed