When two processes on the same Windows machine need to exchange data quickly, few facilities are as direct as shared memory. Rather than serializing a message and handing it to a transport layer, two processes arrange for the very same physical pages of RAM to appear inside both of their virtual address spaces, so that a write performed by one becomes instantly readable by the other. On Windows this arrangement is built on section objects, created and opened through the CreateFileMapping and OpenFileMapping functions and brought into an address space with MapViewOfFile. Understanding how these pieces fit together, how the kernel manages the underlying pages, and how to synchronize access with named mutexes and semaphores is essential for anyone building high throughput pipelines, from video processing graphs to inter-service buffers on a single host.

Why Sharing Physical Pages Matters

Every process on Windows normally owns a private virtual address space. The pages that back a heap allocation in one process are completely invisible to every other process, which is exactly what gives the operating system its isolation guarantees. Interprocess communication mechanisms such as pipes, mailslots, and sockets work around that isolation by copying data, usually twice: once from the sender into a kernel buffer, and once from the kernel buffer into the receiver. For small control messages that cost is trivial, but when a workload moves tens or hundreds of megabytes per second, the repeated copying burns memory bandwidth and CPU time that could be spent doing real work.

Shared memory sidesteps the copy entirely. A section object represents a span of storage, either backed by a file on disk or by the system paging file. When two processes map views of the same section, the memory manager points both view ranges at the same page frames. A memcpy executed by the producer lands in physical memory that the consumer can dereference at its next instruction. The kernel is involved only when views are created or torn down, not on every byte transferred, which is why shared memory remains the fastest general purpose channel for bulk data on one machine.

The speed comes with a responsibility that message based transports handle for you. Because the kernel does not copy, it also does not order, validate, or bracket the transfer. Both parties must agree on a layout inside the shared region, and both must coordinate so that one is not reading bytes the other is halfway through writing. That coordination problem dominates the design of any real shared memory protocol, and it leads directly to synchronization primitives such as named mutexes and semaphores.

Section Objects and the Mapping Functions

A section object is the kernel level anchor for shared storage. CreateFileMapping takes a file handle, security attributes, protection flags, a maximum size, and an optional name, and returns a handle to a section. When the file handle is the invalid handle value, the section is backed by the paging file instead of a real file, producing a region of zero initialized RAM that exists only as long as some process holds a handle open. This page file backed variant is the standard way to build pure memory sharing with no disk component.

Names are the discovery mechanism. A section created with a name such as Local\VideoFrameBuffer enters the object namespace of the current session, while the Global prefix makes it visible across sessions for appropriately privileged processes. A second process that knows the name calls OpenFileMapping with the same string and receives its own handle to the identical underlying object. Handles can also be inherited by child processes or duplicated with DuplicateHandle when a name is undesirable, but naming is by far the most common scheme in pairs of cooperating applications.

No bytes become addressable until a view is mapped. MapViewOfFile takes a section handle, the desired access, a file offset, and a view size, and returns a pointer into the caller's address space. Each process picks its own address, so the same physical byte generally lives at different virtual addresses in producer and consumer, which is why pointers must never be stored inside the shared region; offsets relative to the view base are the portable currency. UnmapViewOfFile releases the view, and closing the last handle to the section lets the kernel reclaim the storage.

Access Protection and Backing Files

Protection flags on CreateFileMapping govern what views may do. PAGE_READWRITE allows read and write views, PAGE_READONLY restricts mapping to read only access, and PAGE_WRITECOPY asks for copy on write semantics, in which a write to a shared file backed page first manufactures a private copy so the modification stays local. MapViewOfFile's FILE_MAP_READ, FILE_MAP_WRITE, and FILE_MAP_ALL_ACCESS requests must be compatible with the protection declared when the section was created, or the call fails with an access error.

Choosing FILE_MAP_COPY copy on write sounds like shared memory but behaves very differently, because after the first write the two processes no longer see the same pages. True bidirectional sharing requires read write views. Copy on write still has legitimate roles, for example loading a large template image that worker processes mostly read but occasionally annotate privately, but pipelines that genuinely exchange frames need PAGE_READWRITE so that physical page identity is preserved for both directions.

The backing choice also carries performance meaning. File backed sections are paged directly against the disk image and are the basis of memory mapped file IO and of the way the loader brings executable images into memory. Paging file backed sections trade that persistence for simplicity, and they participate fully in ordinary memory management: under memory pressure their contents can be pushed to the page file, but while resident they behave like ordinary RAM. For media pipelines and other latency sensitive work, paging file backed sections created by CreateFileMapping with an invalid file handle are the tool of choice.

Synchronization With Named Mutexes and Semaphores

Shared pages provide capacity but no discipline. The classic coordination device is a pair of process shared synchronization objects that live in the same object namespace as the section itself, which is why they too are named. CreateMutex creates a named mutex that any process can open, and the mutex enforces mutual exclusion over a critical stretch of code. A producer that needs exclusive access to a control block inside the shared region waits on the mutex, updates the shared fields, and releases it, guaranteeing that no consumer simultaneously observes a torn, half written structure.

Semaphores solve the counting problem that mutual exclusion does not address. CreateSemaphore establishes a count, and WaitForSingleObject decrements it while ReleaseSemaphore increments it. In a producer consumer ring, one semaphore counts filled slots and another counts empty slots. The producer waits on the empty semaphore, writes a frame, then releases the filled semaphore; the consumer mirrors the pattern in reverse. The kernel blocks and wakes threads as counts change, so no spinning wastes CPU, and because the objects are named, an unrelated consumer process can join the protocol late by opening the objects with OpenSemaphore and OpenMutex.

A few subtleties deserve attention. Wait functions accept timeouts, which prevents a stalled partner from freezing the whole pipeline forever. The state after a WAIT_TIMEOUT must be treated as no work available rather than an error. Abandoned mutexes are another hazard: if a process holding a mutex terminates without releasing it, waiters receive WAIT_ABANDONED and must recognize that the protected data may be inconsistent and repair it before continuing. Robust shared memory designs centralize the wait logic in a small library both sides link against so the edge cases are handled once, correctly.

Race Conditions and Memory Ordering Pitfalls

Even with mutexes and semaphores in place, shared memory invites races that single process code never encounters. A flag written to indicate that a buffer is ready can be observed by the reader before the buffer contents are visible, because the compiler and the CPU may reorder independent stores for performance. On x86 hardware most ordinary loads and stores appear in a strong order, but the compiler alone is free to hoist or sink accesses, and on ARM based Windows systems the hardware itself permits far more reordering. Declaring shared control fields volatile discourages some compiler tricks but is not a complete answer.

The disciplined approach uses interlocked operations and explicit memory semantics. InterlockedExchange and its relatives perform atomic read modify write cycles with full ordering, and the Acquire and Release variants of the interlocked family give precisely the semantics a producer consumer protocol needs. A release write when publishing the ready flag guarantees every earlier store, including the payload writes, is visible before the flag; an acquire read when observing the flag guarantees later loads see the payload. Relying on the synchronization object wait and release calls is usually sufficient because those functions carry acquire and release semantics themselves, but designs that peek at flags without waiting must reason about ordering explicitly.

Torn reads are the opposite failure, where a reader catches a structure mid update. Wide counters that exceed the natural atomic size, variable length payloads, and double word fields on 32 bit builds all invite partial observation. Sequence locking is a popular remedy inside shared regions: the writer increments a sequence number to an odd value, performs the update, and increments again to an even value, while readers retry whenever they see an odd number or a changed pair. Combined with a named semaphore for blocking notification, this scheme lets readers avoid taking the mutex on the fast path, which keeps frame throughput high.

Shared Memory in Media Pipelines

Video and audio processing chains are the canonical large payload workload, and shared memory shows up across the Windows media stack. A capture filter receives frames from a camera driver, must hand them to a decoder, a color conversion stage, and a renderer, and each handoff through a message based transport would copy every frame. With megapixel frames arriving thirty or sixty times per second, those copies measurably raise CPU usage and add latency. Instead, pipelines negotiate an allocator that hands every stage pointers into a pool of shared sections, so a frame written once flows through the graph by reference.

A typical custom pipeline structures a named section as a control header followed by a ring of fixed size frame slots. The header holds the negotiated format, slot count, and indices protected by a named mutex, while filled and empty named semaphores pace the producer and consumers. The producer maps the section once, waits for an empty slot, writes the frame, publishes its index under the mutex, and releases the filled semaphore. A consumer, perhaps an encoder process, performs the mirror image and returns the slot to the empty pool. Adding a second consumer, such as a preview window, is as simple as opening the same named objects.

The same pattern serves screen capture tools, game capture overlays, machine vision stages, and audio mixing, anywhere payload size times rate makes copying the bottleneck. The design lessons generalize: name the section and its synchronization objects in one agreed namespace, never store pointers across process boundaries, treat every control field with explicit ordering, and plan for the abandoned mutex and dead partner cases from the start. Get those right and section objects deliver the closest thing Windows offers to zero copy interchange between otherwise isolated processes, a capability whose simplicity and raw speed explain why it underpins so many performance critical products shipping today.