GeoLocator v1.4: Multi-Task Training, and a Data Leak
Archived — superseded by GeoLocator 4.0
This post is kept for the record. The performance and training figures it originally quoted were never measured on a published benchmark, so they have been removed rather than restated. GeoLocator 4.0 is the current engine and the first version measured on our own published benchmark.
v1.4 was the first real architectural step away from the v1.0 baseline. It kept the idea of classifying a photo into a geographic cluster, but trained on three tasks at once and fixed the clustering scheme that had confused the baseline.
Its results were not valid
The validation split overlapped the training data, so the model was partly being scored on photographs it had already seen. Any accuracy or distance figure from this run says more about the leak than about the model, which is why none is quoted here. The architecture work was still worth keeping; the numbers were not.
Architecture: three heads instead of one
Instead of a single cluster classifier, v1.4 asked the same backbone to answer three related questions. The idea was that country identity and a local offset are easier to learn together than a cluster label alone.
Backbone
ConvNeXt XXLarge
CLIP-pretrained weights
Model type
MultiTaskGeoModel
Three specialised output heads
Output heads
Geo head
Picks the geographic cluster
Country head
Auxiliary country classification
Refinement head
Predicts a latitude/longitude offset inside the cluster
Clustering fix: border-respecting
Unlike v1.0, clusters were built so that they do not straddle international borders. That helps the country head and stops a single label covering two places that look nothing alike.
Loss and optimisation
A custom multi-task loss combined cross-entropy on the cluster head, cross-entropy on the country head and a regression term on the refinement head, with the refinement term weighted most heavily so that sub-cluster precision was not drowned out by the classification terms.
Training setup
- AdamW with a cosine learning-rate schedule.
- Mixed-precision training with gradient scaling.
- Random-crop training transforms, centre crop for validation, colour jitter and coarse dropout.
- Checkpoints kept on the best validation score — which, because of the leak, was not meaningful.
What came next
The architecture survived; the evaluation did not. v2.0 retrained the same multi-task idea on a clean, non-overlapping split so that the results meant something. Today's engine works differently again — it reasons about visual evidence and resolves coordinates from a gazetteer. Try the current engine.