A shared point encoder followed by symmetric pooling gives a scan-level representation that does not change when input points are reordered.
Point-set permutation invariance and pooled features
Treat order as meaningless
A scanner may return the same surface points in a different order because of packet timing or crop serialization. A scan-level classifier should not change its prediction under that permutation. Apply the same feature network to every point, then aggregate with a symmetric operation such as maximum or mean. The code verifies the pooled vector is unchanged after an explicit permutation. The choice of pooling still changes sensitivity to rare defects and noise.
Know what max pooling sees
Elementwise maximum keeps the strongest activation for each feature channel. One strongly activated point can represent a small damaged edge, but a spurious reflection can also dominate. Mean pooling uses information from many points and may dilute a tiny fault. Compare both on defect size and missing-return slices. A simple global representation does not preserve detailed neighborhood geometry; if local shape matters, add measured local grouping only after checking a set baseline.
Keep per-point and global tasks separate
Pallet classification yields one label per scan. Segmenting damaged surface points requires outputs tied to the individual input points. In that case combine each point feature with a broadcast global feature before prediction, and restore predictions to the raw scan IDs after sampling. Sorting or shuffling input points should permute per-point outputs in the same way rather than leave them fixed. Segmentation alignment has a parallel issue in images.
Test missing density and transforms
Permutation invariance alone does not guarantee resilience to point deletion, changed sensor density, translation, rotation or scale. Evaluate deliberate resampling and coordinate perturbations within the sensor tolerance, then inspect whether the model uses the pallet rather than its scan rig. Group held-out scans by site and calibration period. The coordinate contract fixes what each number means before any invariant architecture is useful.
Measure the inference budget
A shared encoder processes every retained point; doubling the point budget roughly doubles pointwise work and activation storage for a fixed width. Pooling itself is cheap. Track p95 latency for raw scan decode, outlier rejection, sampling and network inference. A tiny network forward can be operationally slow when preprocessing copies millions of coordinates. The project gates the complete inspection decision.
Implementation
import torch
from torch import nn
torch.manual_seed(47)
pallet_points = torch.tensor([[[0.1, 0.2, 0.0], [0.3, 0.1, 0.1],
[0.2, 0.4, 0.2], [0.4, 0.3, 0.0]]])
point_encoder = nn.Sequential(nn.Linear(3, 16), nn.ReLU(),
nn.Linear(16, 8))
point_features = point_encoder(pallet_points)
scan_features = point_features.amax(dim=1)
reordered_features = point_encoder(pallet_points[:, [2, 0, 3, 1], :]).amax(dim=1)
assert scan_features.shape == (1, 8)
assert torch.allclose(scan_features, reordered_features)
inspection_head = nn.Linear(8, 2)
inspection_logits = inspection_head(scan_features)
assert inspection_logits.shape == (1, 2)Performance and operating cost
For B scans, N points and a point encoder with F operations per point, pointwise work is O(BNF) and activations scale with O(BNW) for hidden width W. Symmetric pooling adds O(BNW) comparisons or additions and a small O(BW) output. Local-neighbor architectures can incur extra search and grouping costs; a global set encoder avoids those costs but may miss fine spatial structure. This synthetic shape and reorder check proves permutation behavior only, not inspection accuracy.
Common Mistakes
- Do not mistake order invariance for rotation or scale invariance.
- Do not use one global class logit as if it located damaged points.
- Do not hide reflective outliers behind a mean accuracy score.
