GeoLocator v2.0: Multi-Task Training on a Clean Split
Archived — superseded by GeoLocator 4.0
This post is kept for the record. The performance, training and "theoretical limit" figures it originally quoted were never measured on a published benchmark, so they have been removed rather than restated. GeoLocator 4.0 is the current engine and the first version measured on our own published benchmark.
v2.0 exists because of a mistake. v1.4 had trained on data that overlapped its own validation set, so its results could not be trusted. v2.0 kept the multi-task architecture and rebuilt the part that actually mattered: the split.
A validation split you can believe
Training and validation photographs were separated so that neither set could contain the other's images. Honest evaluation made the model look worse — and that was the point: it was the first version whose weaknesses were visible at all.
What it showed was sobering. A cluster classifier trained this way could place a photo in roughly the right part of the world, but it was nowhere near town-level accuracy, and the gap between training and validation loss made it clear how much of v1.4's apparent skill had been memorisation.
Architecture
The same multi-task framework as v1.4: one backbone, three heads, trained together.
Backbone
ConvNeXt XXLarge
CLIP-pretrained weights
Optimisation
torch.compile
Faster training and inference
Multi-task output heads
Geo head
Coarse location as a cluster classification
Country head
Auxiliary country classification for spatial awareness
Refinement head
Latitude/longitude offset for sub-cluster precision
Training setup
- Distributed data-parallel training with synchronised batch norm.
- Gradient checkpointing to fit the larger backbone in memory.
- Mixed precision with gradient scaling.
- AdamW with a cosine learning-rate schedule and gradient clipping.
- Border-respecting clusters carried over from v1.4.
The loss function
A custom multi-task loss: label-smoothed cross-entropy on the cluster head, cross-entropy on the country head and a weighted regression term on the refinement head, with the refinement term weighted most heavily so that the model kept caring about precision inside a cluster.
What we took from it
v2.0 was the first version we could reason about honestly, and it convinced us that cluster classification had a ceiling: the label space itself limits how precise the answer can be, however good the backbone is. That pushed the project towards reading evidence and resolving places by name, which is what v3.0 started and GeoLocator 4.0 finished. Try the current engine.