GeoLocator v1.0: The First Baseline
Archived — superseded by GeoLocator 4.0
This post is kept for the record. The performance and training figures it originally quoted were never measured on a published benchmark, so they have been removed rather than restated. GeoLocator 4.0 is the current engine and the first version measured on our own published benchmark.
v1.0 was where the project started: a single-head classifier that sorted a photograph into one of a set of geographic clusters and returned the centre of the winning cluster as its guess. It was a baseline in the truest sense — something to beat, not something to ship.
It did not work well
The baseline regularly failed to identify even the right country. Two causes stood out: the training images were scraped from a general photo dataset and many of them carried no useful geographic signal at all, and the clusters themselves were drawn without regard for borders.
What v1.0 was
A convolutional backbone produced image features, one classification head chose a cluster, and the predicted coordinates were simply that cluster's centre. Evaluation used Haversine distance between the predicted point and the true one, and checkpoints were kept on validation loss.
Backbone
ConvNeXt Tiny
Convolutional backbone
Architecture
Single head
Cluster classifier
Loss
CrossEntropy
Standard classification
Training setup
- AdamW with a cosine learning-rate schedule and gradient clipping.
- Mixed-precision training.
- Early stopping on validation loss, with the best checkpoint kept.
- Standard augmentation: random crop, horizontal flip, colour jitter and coarse dropout.
What went wrong
The training images came from a general-purpose photo dataset. Plenty of them were indoor shots, close-ups or artwork — pictures with a coordinate attached but nothing in the frame that could ever identify it. A model trained on those learns noise.
Clustering across borders
Clusters were allowed to span international borders, so a single label could cover two countries with completely different signage, road markings and architecture. That confuses training and makes the output hard to interpret. v1.4 fixed it with border-respecting clusters.
What it taught us
Two lessons carried through every later version: the data matters more than the backbone, and the label space has to match the way the world actually looks. Those pushed us towards border-respecting clusters and multi-task training in v1.4, and eventually away from cluster classification altogether — today's engine reasons about visual evidence and resolves coordinates from a gazetteer instead. Try the current engine.