A sequence convolution is causal only when its padding and indexing prevent future samples from influencing an earlier output.
Causal dilated convolution and receptive field
State the forecast clock
For a warehouse queue at minute 47, a forecast must use measurements observable at or before minute 47. A one-dimensional convolution with symmetric padding can place future values in the window around an output even if its output shape looks right. Use left padding of dilation times kernel-width-minus-one for stride-one causal convolution, then verify by perturbing a later input and checking every earlier output. The code does that directly. Recurrent state has a different but related time-boundary problem.
Calculate the receptive field
For a chain of stride-one causal convolutions with kernel width k and dilation values d, the receptive-field length is one plus the sum of (k minus one) times each d. Three width-three layers with dilations one, two and four see fifteen time positions. A lookback shorter than this can leave much of the nominal receptive field filled with padding. Receptive field is a maximum context span, not evidence that every old measurement matters equally. Record all layer widths, dilations, strides and crop rules.
Keep feature latency explicit
Event time and availability time differ. A database metric observed at 47 might not be available until 49. If the forecast runs at 47, including it violates the information boundary even though its event timestamp is not in the future. Build training windows from the as-available feature view, not a backfilled table. For missing minutes, choose a gap marker, forward-fill rule or explicit mask and apply it in serving. Time-safe graph splits rely on the same observable-at principle.
Compare long context with a short baseline
Dilations expand coverage without proportionally expanding kernel parameters, but a long nominal lookback can still add compute and noise. Compare with last-value, seasonal and short-window models on identical rolling origins. Report errors by demand regime, missingness and holiday or incident periods. If the long-context model helps only on windows after a future-data backfill, the gain is suspect. The project uses a rolling release gate.
Check deployment alignment
A causal model trained to output a value at every input time must select the final output for a next-horizon forecast and must shift labels consistently. An off-by-one label can yield attractive validation loss while serving a forecast for the already observed interval. Save an input window with timestamps, expected horizon and model output, and replay it after changes. Do not treat matching tensor lengths as proof of correct forecast alignment.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
queue_history = torch.arange(11, dtype=torch.float32).view(1, 1, 11)
causal_filter = nn.Conv1d(1, 1, kernel_size=3, dilation=2, bias=False)
left_padding = 2 * (3 - 1)
baseline = causal_filter(functional.pad(queue_history, (left_padding, 0)))
changed_future = queue_history.clone()
changed_future[:, :, -1] = 900
replayed = causal_filter(functional.pad(changed_future, (left_padding, 0)))
assert baseline.shape == queue_history.shape
torch.testing.assert_close(baseline[:, :, :-1], replayed[:, :, :-1])
assert 1 + (3 - 1) * sum((1, 2, 4)) == 15Performance and operating cost
For B windows, input channels C_in, output channels C_out, length T and kernel width k, a direct convolution is on the order of O(BTC_inC_outk) operations per layer; dilation changes spacing, not the number of kernel weights. Activations cost O(BTC_out) per layer in training, apart from backend buffers. Stacking layers increases the receptive field and stored activations. The test proves causal indexing for one convolution but does not measure forecast skill or data-availability leakage.
Common Mistakes
- Do not use symmetric same-padding and assume it is causal.
- Do not train on backfilled features unavailable at prediction time.
- Do not confuse output alignment with the forecast horizon.
Read next
- Quantile forecast loss and crossing policy
- Project: forecast warehouse queues with a causal temporal model
- Recurrent hidden-state resets and truncated gradients
- Causal target shifts and padding-aware loss
- Inference contracts: preserve preprocessing and measure tail latency
Continue the workflow: Video temporal convolutions and clip aggregation.
Continue the workflow: DQN transitions, replay sampling and reward contracts.
