The stem output \(F_{\text {stem}} \in \mathbb {R}^{H/2 \times W/2 \times 32}\) preserves sufficient spatial detail for subsequent stages to learn local patterns.
\end{aligned}$$ (8) where s is the stride of the depthwise convolution, \(\oplus\) denotes element-wise addition, and the shortcut is omitted when dimensions differ.
Two stride-2 blocks reduce spatial size by a factor of four, yielding \(F_{1,\text {out}} \in \mathbb {R}^{H/8 \times W/8 \times 16}\).
A stride-2 block reaches \(H/32 \times W/32\), and three stride-1 blocks deepen the representation.
The feature pyramid runs from \(H/2 \times W/2\) to \(H/32 \times W/32\), providing a multi-scale representation suited to polyps of varying size, and is optimized for speed through depthwise convolutions and careful channel management.