Spatial behavior in CNNs is controlled by convolution, stride, padding, and feature maps.
This guide explains how these hyperparameters shape model outputs and why they matter in real architectures.
Cross-posted from Zeromath. Original article: https://zeromathai.com/en/pooling-activation-layers-en/
Why Spatial Understanding > Just "Layers"
Most explanations of CNNs focus on stacking layers.
But in practice, what really matters is:
- how spatial dimensions change
- how features are extracted locally
- how resolution evolves
CNNs are not just computation pipelines.
They are spatial transformation systems.
Convolution = Sliding Pattern Matching
A convolution filter:
- slides over the image
- computes weighted sums
- produces a feature map
Key idea:
Same detector applied everywhere
This gives:
- translation invariance
- parameter efficiency
Filters and Feature Maps
Example:
- Input: 6 × 6 × 3
- Filter: 3 × 3 × 3
Output:
- 4 × 4 feature map
Multiple filters:
- 32 filters → 32 feature maps
So output becomes: