Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Training memory profiling and allocator boundaries

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A memory report must distinguish live tensors, allocator reservations, optimizer initialization and the peak of the actual forward–backward–step path.

Choose the measured boundary

State whether the report covers data transfer, forward, backward, gradient clipping, optimizer step or all of them. Adam-like optimizers may allocate moment tensors at the first step, so a warm forward-only profile undercounts a real training update. Measure a representative full update after optimizer state exists. Reset peak statistics at the chosen boundary, synchronize the accelerator before and after timing, and log batch shape, dtype, model revision and gradient-accumulation count. An effective batch may include several microbatches.

Separate allocated and reserved bytes

Live tensor allocation reports memory currently occupied by tensors; an allocator may reserve extra blocks for reuse. Peak allocated and peak reserved therefore answer different questions. A process can hit a device limit even if a later snapshot shows fewer live tensors. Clearing unused cached blocks does not free memory still held by model weights or activations, and calling a cache-clear operation after every step can reduce throughput. Collect both metrics and compare them with device-level process usage when diagnosing an out-of-memory event.

Account for hidden retained graphs

Appending loss tensors or outputs that still reference a computation graph to a Python list can keep activations alive across steps. Log scalar values after detaching rather than storing graphs for dashboards. Forward hooks and explanation maps can cause similar retention. A rising allocated baseline over repeated identical steps deserves an object-lifetime audit before changing batch size. The profiler code captures one representative step but also needs a repeated-step trend in production. Diagnostic probes should not stay attached to the training graph.

Compare methods under identical work

A fair activation-checkpointing or mixed-precision comparison holds model, microbatch, accepted examples, optimizer state and input size fixed. Then report peak allocated, peak reserved and synchronized update time. Only after that parity check should a larger batch be tested as a separate throughput experiment. Batch size can change normalization statistics, gradient noise and data order. Recomputation trades time for memory; it does not automatically improve examples per second.

Keep device assumptions explicit

CUDA counters describe CUDA allocator behavior and do not measure CPU RAM, pinned transfer buffers or every external library allocation. On a CPU-only runner, return an unavailable marker for those GPU counters rather than reporting zero as if no memory were used. Record device, driver/runtime revision and compilation mode. The applied project requires measured benefit on a target runner and explains when the technique is worth deploying.

Implementation

python
import time
import torch
from torch import nn
from torch.nn import functional as functional

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
receipt_model = nn.Sequential(nn.Linear(16, 32), nn.ReLU(), nn.Linear(32, 3)).to(device)
optimizer = torch.optim.AdamW(receipt_model.parameters(), lr=0.0007)
features = torch.randn(6, 16, device=device)
labels = torch.tensor([0, 2, 1, 2, 0, 1], device=device)

def training_update():
    optimizer.zero_grad(set_to_none=True)
    loss = functional.cross_entropy(receipt_model(features), labels)
    loss.backward()
    optimizer.step()
    return float(loss.detach())

training_update()  # initialize optimizer state before the measured step
if device.type == "cuda":
    torch.cuda.synchronize()
    torch.cuda.reset_peak_memory_stats()
started = time.perf_counter()
observed_loss = training_update()
if device.type == "cuda":
    torch.cuda.synchronize()
elapsed_seconds = time.perf_counter() - started
peak_allocated = torch.cuda.max_memory_allocated() if device.type == "cuda" else None
peak_reserved = torch.cuda.max_memory_reserved() if device.type == "cuda" else None
assert elapsed_seconds >= 0 and observed_loss >= 0
assert peak_allocated is None or peak_reserved >= peak_allocated

Performance and operating cost

Profiling adds small counter and synchronization overhead, but synchronization changes asynchronous execution timing and should be applied consistently across compared runs. A training step is dominated by model work, not counter reads. Peak memory depends on model parameters, gradients, optimizer moments, activations, batch inputs and allocator behavior. The code intentionally returns no GPU-memory number on CPU; it would be false precision to substitute process RAM or zero. Repeat warm measurements and report the distribution, not one favorable run.

Common Mistakes

  • Do not compare a forward-only baseline with a full checkpointed optimizer update.
  • Do not treat allocator-reserved bytes as identical to live-tensor allocation.
  • Do not report zero GPU memory on a CPU-only machine as a device benchmark.

Read next

ai-data
deep-learning
Storage details