GOOSE-M2F: Adapting Mask2Former for High-Fidelity, Long-Tailed Fine-Grained Semantic Segmentation in Unstructured Outdoor Terrain
64 fine-grained terrain classes, some occupying fewer than 50 pixels an image. Our Mask2Former adaptation placed 3rd on the GOOSE 2D leaderboard at 70.08% composite mIoU.
Jyothiraditya Lingam1, Nikhileswara Rao Sulake1, Sai Manikanta Eswar Machara1
1 Department of Computer Science and Engineering, Rajiv Gandhi University of Knowledge Technologies, Nuzvid, India
GOOSE 2D FGSS Challenge, ICRA 2026 · Technical Report

A benchmark that refuses to be easy
The GOOSE 2D Fine-Grained Semantic Segmentation Challenge asks a model to label outdoor terrain into 64 separate classes, not the dozen or so categories most segmentation benchmarks use. Unstructured outdoor scenes make this harder still: trails, vegetation, and ground cover blend into each other, and a meaningful share of the 64 classes appear so rarely that some occupy fewer than 50 pixels in an entire image. Standard segmentation training tends to quietly abandon classes like these, since a network can get most of its loss down just by getting the big, common regions right.
Extending Mask2Former for the long tail
We built on Mask2Former with a Swin-Large backbone and made three targeted changes rather than a full redesign. We expanded the number of object queries to 200, since the default query count saturates well before covering 64 distinct classes across a scene. We added a Feature Refinement Module that combines dilated convolutions with channel and spatial attention, giving the model more room to sharpen features for small, easily missed structures before they reach the transformer decoder. And we added an auxiliary per-pixel supervision head that runs only during training, feeding direct gradients back for rare classes that would otherwise get drowned out by the mask-level loss.
Training and inference details that mattered
Architecture changes alone were not enough. We paired them with distribution-balanced loss weighting, rare-class copy-paste augmentation, dynamic IoU-aware re-weighting during training, and an exponential moving average of model weights. At inference time, a sliding-window pass with 2D Gaussian kernel blending and four-scale test-time augmentation added another 10.57 percentage points of composite mIoU on its own, one of the larger single contributions in the whole pipeline.
Where GOOSE-M2F landed
The final system reached 70.08 percent official composite mIoU on the GOOSE 2D FGSS leaderboard, splitting into 63.55 percent on fine-grained classes and 76.61 percent on coarse ones, and placed third overall on the challenge. We wrote the approach up as a technical report and released the code and trained weights publicly, since a lot of the value in a long-tailed benchmark like this comes from other teams being able to build on what worked and what did not.

A note on the team
This project was a genuinely small team effort with Nikhileswara Rao Sulake and Jyothiraditya Lingam, who was still in his first year of engineering while this work was happening. Working alongside someone that early in their degree taking a challenge this seriously was, honestly, one of the more encouraging parts of the project.