GeoLocator v4.0: Eyes, Judgment, and a Map — Our New Three-Stage Engine
Correction 2 of 2
Update, September 23, 2026: a new, larger benchmark — and a median distance error that got worse on purpose
This post has now been corrected twice. Both notes stay on the page, in order, because a page that quietly rewrites its own numbers is not worth reading.
On September 22 we replaced the engine's vision stage and re-measured everything on a new benchmark: 478 outdoor photographs instead of 69, across 44 countries, every one of them captured and first published after any model in the pipeline finished training. Two headline figures moved in opposite directions, and only one of those movements is the engine changing.
Country accuracy went up: 83% → 91.2%. That is measured on a set roughly seven times larger and harder than the old one, and the 95% interval narrowed from ±8.3 points to ±2.7. It is a genuine improvement, and it is now measured well enough to state.
Median distance error went up: 1.8 km → 36.7 km. That is not the engine getting worse. The old 69 photographs were selected as photographs of towns, close to town centres, and the engine answers by pinning to the centre of the town it names — so on a set of town photographs the pin was nearly always right and the number flattered us. The new set is ordinary outdoor photography: coastline, farmland, forest, roads. On the photographs in the new set taken within 1 km of a town, the median error is 2.8 km — the honest descendant of the old 1.8 km. The 36.7 km figure is what happens when you stop measuring only towns, and it is the number we lead with rather than the flattering subset.
Every figure in the body of this post now comes from that single run. The older note below keeps the figures it was written with, so you can see exactly what changed. The old set has been retired; the reasons are in why the benchmark changed.
Correction 1 of 2 · kept for the record · figures superseded by the note above
Update, September 22, 2026: re-measured on the deployed engine
We re-ran the same 69 benchmark photos through the engine we deploy, exactly as the site runs it. It names the right country 83% of the time (57 of 69), with a 1.8 km median distance error; 36% of results land within 1 km, 61% within 25 km, 68% within 200 km, 83% within 750 km and 93% within 2,500 km. An analysis takes about 25 seconds.
The figures first published here (84% country accuracy, a 2 km median error, 35% within 1 km, about 15 seconds) were measured on our development pipeline, which also gave the decision steps a second model's written description of each photo. The engine we deploy does not use that description, so this post now reports the deployed engine. Two claims we could not re-measure on it, the vision model's score on its own and a confidence comparison from an earlier benchmark, have been removed; the confidence figures below are recomputed from the new run.
Every photo carries evidence. A licence plate. The wording on a stop sign. Which side of the road the cars are on. The shape of a bollard, the wiring on a utility pole, the roofline of a house. v3.0 taught us that reasoning over that evidence matters. v4.0 rebuilds the entire pipeline around that lesson: we split the job into three specialists — one that sees, one that decides, and one that turns a decision into a pin on the map.
v4.0 is the engine behind every GeoLocator photo geolocation analysis. On our current benchmark — 478 outdoor photographs from Wikimedia Commons, every one of them taken and published after training ended — it names the right country 91.2% of the time, the median analysis takes 8.5 seconds, and the median distance error is 36.7 km. The demo and accounts are free.
That single accuracy figure is also the least useful thing on this page. The same engine gets the country right 99.2% of the time on a photograph with a readable place name somewhere in the frame, and 84.4% of the time on one with no readable text at all — with median errors of 10.0 km and 125.8 km respectively. A street scene and a photograph of a forest are not the same problem, and the old version of this post averaged them into one number. This one stops.
Key facts
GeoLocator v4.0 at a glance
- Released: September 21, 2026. Replaces the v3.0 engine as the production engine. Vision stage replaced and everything re-measured on September 22, 2026.
- Architecture: three stages — a frontier vision model (eyes), a System One decision model (judgment), and a GeoNames gazetteer resolver (map).
- Benchmark: 478 outdoor photographs from Wikimedia Commons, 44 countries, all captured and first published on or after March 1, 2026. Two analyses errored and are counted as misses.
- Country accuracy: 91.2% traffic-weighted, 95% interval 88.6–93.9%. Raw pooled: 91.4%, 437 of 478 photographs (88.6–93.6%).
- Distance: 36.7 km median error (37.6 km raw). 11.1% of results land within 1 km, 25.5% within 5 km, 45.2% within 25 km, 74.9% within 200 km, 93.0% within 750 km, 98.3% within 2,500 km.
- What that depends on — readable text: 99.2% country accuracy and a 10.0 km median error with place-identifying text in the frame (122 photos); 95.1% and 20.6 km with other text (144); 84.4% and 125.8 km with no readable text at all (212).
- What that depends on — how remote the spot is: within 1 km of a known town, 96.9% and a 2.8 km median error (127 photos); more than 25 km from one, 87.7% and 174.4 km (65).
- Knowledge base: 2,265 curated GeoGuessr-style meta clues across 136 countries, including explicit "how to tell country X from country Y" rules.
- Resolver: GeoNames gazetteer with 34,000+ cities — region, then city, then coordinates.
- Confidence: answers scoring 80 or above had the right country 96.7% of the time (408 of 422); below 80, 53.7% (29 of 54).
- Speed and cost: 8.5 s median, 11.4 s at the 90th percentile, 14.4 s at the 95th, 26.4 s for the slowest photo in the set. $0.0040 per photo in model cost — the whole 478-photo run cost $1.90 and took 23 minutes. The service is free to use.
- Scope: outdoor photographs only. Indoor shots, food, portraits and close-ups were excluded from the benchmark by design, so this is accuracy on outdoor photographs, not on arbitrary uploads.
What actually drives accuracy
Two things dominate everything else, and neither of them is the engine. The first is whether there is readable text in the photograph. The second is how close the photographer was standing to a place the gazetteer has heard of. Once you split the set on either of those, the single averaged number stops being a description of anything.
Readable text is the single strongest signal
| Text in image | Photos | Country | Median error |
|---|---|---|---|
| Place-identifying text | 122 | 99.2% | 10.0 km |
| Text, but not a place name | 144 | 95.1% | 20.6 km |
| No readable text at all | 212 | 84.4% | 125.8 km |
A shop sign, a road sign, a bus destination board: one legible place name and the problem collapses into a lookup. Strip that away and the engine is back to reasoning from vegetation, light, road furniture and building style — which still gets the country right more than four times in five, but puts the pin an order of magnitude further out. Nearly half of the benchmark (212 of 478 photographs) is in that hardest class, which is deliberate.
And distance to the nearest town the map knows
| Distance to nearest gazetteer town | Photos | Country | Median error |
|---|---|---|---|
| Under 1 km | 127 | 96.9% | 2.8 km |
| 1–5 km | 156 | 91.7% | 9.8 km |
| 5–25 km | 130 | 87.7% | 73.9 km |
| Over 25 km | 65 | 87.7% | 174.4 km |
This is a property of how the engine answers, and it is worth being blunt about. Stage 3 does not regress a latitude and longitude out of pixels; it names a town and returns that town's centre. On a photograph taken inside a town, that is an excellent answer. On a photograph taken 40 km up a valley, the best possible answer the design can produce is still the valley's nearest town — so the distance error has a floor that has nothing to do with how well the photo was read. Country accuracy barely moves across these rows — 96.9% at the top, 87.7% at the bottom. The distance column goes from 2.8 km to 174.4 km.
The same story, by what the photograph is of
| Photo type | Photos | Country | Median error |
|---|---|---|---|
| Landmark | 44 | 100.0% | 8.8 km |
| Coast | 26 | 100.0% | 41.3 km |
| Night | 15 | 100.0% | 31.3 km |
| Village | 12 | 100.0% | 50.6 km |
| Street | 81 | 97.5% | 5.6 km |
| Building in context | 138 | 94.9% | 25.3 km |
| Square | 20 | 90.0% | 26.2 km |
| Road | 18 | 88.9% | 132.9 km |
| Landscape | 83 | 81.9% | 88.1 km |
| Nature | 41 | 68.3% | 241.4 km |
Read the small rows with the caution they deserve: several of the perfect scores rest on a dozen or two photographs, and a single miss would move them a long way. The rows to trust are the large ones — street (81), building in context (138), landscape (83), nature (41) — and they say the same thing as the text table from a different angle. A street scene is close to solved. A photograph of a forest is a research problem, and the engine gets the country wrong on nearly a third of them.
Country accuracy
What is in the frame decides the answer
All 478 photographs, split by photo type and by the text visible in the shot.
under 70% / 70 to 85% / 85 to 95% / 95 to 99.9% / 100%
Hover a cell, or focus the grid and use the arrow keys, for its exact figures.
Place-name text
122 · 99.2%
Other text
144 · 95.1%
No text
212 · 84.4%
Landmark
44 · 100.0%
Coast
26 · 100.0%
Night
15 · 100.0%
Village
12 · 100.0%
Street
81 · 97.5%
Building in context
138 · 94.9%
Square
20 · 90.0%
Road
18 · 88.9%
Landscape
83 · 81.9%
Nature
41 · 68.3%
| Photo type | All text classes | place-identifying text | non-place text | no text at all |
|---|---|---|---|---|
| Landmark | 100.0% from 44 photos | 100.0% from 12 photos | 100.0% from 13 photos | 100.0% from 19 photos |
| Coast | 100.0% from 26 photos | 100.0% from 4 photos (too few to read as a rate) | 100.0% from 4 photos (too few to read as a rate) | 100.0% from 18 photos |
| Night | 100.0% from 15 photos | 100.0% from 6 photos (too few to read as a rate) | 100.0% from 6 photos (too few to read as a rate) | 100.0% from 3 photos (too few to read as a rate) |
| Village | 100.0% from 12 photos | 100.0% from 3 photos (too few to read as a rate) | 100.0% from 2 photos (too few to read as a rate) | 100.0% from 7 photos (too few to read as a rate) |
| Street | 97.5% from 81 photos | 96.7% from 30 photos | 97.7% from 44 photos | 100.0% from 7 photos (too few to read as a rate) |
| Building in context | 94.9% from 138 photos | 100.0% from 45 photos | 97.8% from 46 photos | 87.2% from 47 photos |
| Square | 90.0% from 20 photos | 100.0% from 6 photos (too few to read as a rate) | 71.4% from 7 photos (too few to read as a rate) | 100.0% from 7 photos (too few to read as a rate) |
| Road | 88.9% from 18 photos | 100.0% from 6 photos (too few to read as a rate) | 100.0% from 7 photos (too few to read as a rate) | 60.0% from 5 photos (too few to read as a rate) |
| Landscape | 81.9% from 83 photos | 100.0% from 8 photos (too few to read as a rate) | 75.0% from 8 photos (too few to read as a rate) | 80.6% from 67 photos |
| Nature | 68.3% from 41 photos | 100.0% from 2 photos (too few to read as a rate) | 85.7% from 7 photos (too few to read as a rate) | 62.5% from 32 photos |
| All photo types | 91.4% from 478 photos | 99.2% from 122 photos | 95.1% from 144 photos | 84.4% from 212 photos |
The same curve · split by what the photograph shows
Readable text is the whole story
Three panels, one shared scale. Brighter means more place information in the frame. Hover anywhere to compare all three at the same distance.
Place-identifying text
n = 122 · 99.2% country
median 10.0 km
Non-place text
n = 144 · 95.1% country
median 20.6 km
No text at all
n = 212 · 84.4% country
median 125.8 km
| Text in the photograph | Photographs | Country accuracy | Median error | Within 10 km | Within 100 km | Within 1,000 km |
|---|---|---|---|---|---|---|
| Place-identifying text | 122 | 99.2% | 10.0 km | 50.82% | 89.34% | 98.36% |
| Non-place text | 144 | 95.1% | 20.6 km | 40.97% | 65.97% | 93.06% |
| No text at all | 212 | 84.4% | 125.8 km | 20.28% | 46.23% | 90.09% |
Architecture: Eyes → Judgment → Map
v4.0 splits the work into three stages, gives each job to the component best suited to it, and passes structured evidence between them. Seeing is a perception problem. Deciding between look-alike countries is a knowledge problem. Turning a country and a town name into coordinates is a lookup problem. Three problems, three tools.
Stage 1 · Eyes
Frontier vision model
The vision stage reads the photo and produces ranked country candidates, each with written evidence, plus a nearest-town estimate. It is the perception layer — everything downstream works from what it sees.
Stage 2 · Judgment
System One decision model
Our decision layer disambiguates among the top candidates using a curated knowledge base of 2,265 GeoGuessr-style meta clues across 136 countries — licence plates, scripts and stop-sign wording, driving side, bollards, utility poles, architecture — and gives every answer a confidence score.
Stage 3 · Map
Gazetteer resolver
A GeoNames gazetteer of 34,000+ cities resolves the region, then the city, and returns coordinates. The map step is a lookup rather than a guess — and lookups do not hallucinate.
The knowledge base is the part we are proudest of. It is not a pile of facts; it is organised around the questions that actually decide a photo. When the vision model narrows a photo down to two neighbouring countries that share a script, a road style, and a climate, the decision layer does not need to know everything about both — it needs the handful of rules that tell them apart. Those explicit tell-apart rules are what the 2,265 clues encode, and they are what the decision layer reaches for when the perception model that feeds it cannot separate two candidates.
Why a decision layer, not just a bigger model
The obvious move in 2026 is to throw a larger model at the problem. v4.0 puts a decision layer and its knowledge base on top of the vision model instead, built for one kind of case: photos where the top candidates look alike and only a rule can separate them.
The most useful part does not show up in an accuracy number. The decision layer reports a confidence with every answer. It is a score, not a probability, and on the deployed engine it tracks correctness. Across the 478 photographs, when the confidence was 80 or higher (422 photographs), it had the country right 96.7% of the time — 408 of 422. When it was below 80 (54 photographs), it was right 53.7% of the time — 29 of 54, near enough a coin flip. That gap is the whole point, and it is the one claim on this page that survived a change of vision model and a sevenfold increase in the size of the set. A confidence score that actually tracks correctness is a trust signal you can build on: lean on the high-confidence results, and route the low-confidence ones to a second pass or a human analyst. A bigger model gives you a better guess. A decision layer tells you when to believe it.
Confidence below 80
53.7%
country right, 29 of 54 photographs
Confidence 80 or more
96.7%
country right, 408 of 422 photographs
The buckets underneath tell you how much of that is real. 402 of the 478 answers sit in the top 90–100 bucket and are right 97.0% of the time; 80–90 holds 20 photographs at 90.0%; 70–80 holds 20 at 60.0%; 50–70 holds 32 at 50.0%; and the bottom bucket holds two photographs, which tells you nothing at all. The engine is confident most of the time, and the interesting behaviour lives in the thin middle where it is not.
Confidence calibration
What the score is worth
Confidence below 80
53.7%
right country, 29 of 54
Confidence 80 and above
96.7%
right country, 408 of 422
| Confidence bucket | Photographs | Country correct |
|---|---|---|
| Under 50 | 2 | 50.0% |
| 50-70 | 32 | 50.0% |
| 70-80 | 20 | 60.0% |
| 80-90 | 20 | 90.0% |
| 90-100 | 402 | 97.0% |
| 80 and above (all) | 422 | 96.7% |
| Below 80 (all) | 54 | 53.7% |
Benchmarks
Country accuracy
91.2%
traffic-weighted, 478 outdoor photos · 88.6–93.9%
Median error
36.7 km
across all photo types · 2.8 km inside a town
Within 25 km
45.2%
478 outdoor photos
Within 1 km
11.1%
478 outdoor photos
The headline is traffic-weighted; the raw pooled figure is 91.4% — 437 of 478 photographs, 95% interval 88.6–93.6% — and the median error on the raw basis is 37.6 km. Weighting moves the headline by two tenths of a point, which is the point of reporting both. Two analyses errored outright and are counted as misses rather than dropped. And with 478 photographs rather than 69, no single photograph can move the headline the way one could before — which is most of why the set was rebuilt.
GeoLocator 4.0 · bench40 · 478 outdoor photographs
How often the pin lands within a given distance
Read it as “within x km, the pin is right y% of the time.” Marked points are the six published thresholds. Hover or use the arrow keys to read any other distance.
| Within | Share of photographs | Photographs |
|---|---|---|
| 1 km | 11.1% | 53 of 478 |
| 5 km | 25.5% | 122 of 478 |
| 25 km | 45.2% | 215 of 478 |
| 200 km | 74.9% | 354 of 478 |
| 750 km | 93.0% | 440 of 478 |
| 2,500 km | 98.3% | 468 of 478 |
| Median error | 36.7 km | 478 photographs, 44 countries |
By region, nothing separates cleanly. Europe is 87.7% on 162 photographs (81.7–91.9%), North America 94.9% on 138 (89.9–97.5%), Asia 93.5% on 77 (85.7–97.2%) and the combined rest of the world 91.1% on 101 (83.9–95.2%). Every one of those intervals overlaps the 91.2% headline, so this set cannot support a claim that any one region is genuinely harder than another — only that Europe, which is both the largest stratum and the lowest point estimate, is the most likely candidate if a difference exists.
Country accuracy by region
Every estimate is a range, not a number
Europe
n = 162
North America
n = 138
Asia
n = 77
Rest of world*
n = 101
All photographs
n = 478
87.7%
81.7-91.9
94.9%
89.9-97.5
93.5%
85.7-97.2
91.1%
83.9-95.2
91.2%
88.6-93.9
* Rest of world is Africa, South America and Oceania reported together. Separately they are n = 26, n = 48 and n = 27 - too small to publish on their own, so a single combined figure is the honest one.
| Region | Photographs | Country accuracy | 95% interval |
|---|---|---|---|
| Europe | 162 | 87.7% | 81.7-91.9% |
| North America | 138 | 94.9% | 89.9-97.5% |
| Asia | 77 | 93.5% | 85.7-97.2% |
| Rest of world | 101 | 91.1% | 83.9-95.2% |
| All photographs (traffic-weighted headline) | 478 | 91.2% | 88.6-93.9% |
Where the 478 photographs are
| Miss | Taken | Placed |
|---|---|---|
| 1 | 8.07°N 74.79°W | 10.63°N 85.44°W |
| 2 | 36.35°N 6.17°W | 37.02°N 7.93°W |
| 3 | 50.08°N 14.36°E | 52.23°N 21.01°E |
| 4 | 60.70°N 15.01°E | 62.24°N 25.72°E |
| 5 | 42.90°N 79.86°W | 39.96°N 83.00°W |
| 6 | 51.54°N 0.38°W | 37.77°N 122.42°W |
| 7 | 59.37°N 17.99°E | 54.69°N 25.28°E |
| 8 | 4.58°N 74.30°W | 6.23°S 77.87°W |
| 9 | 14.54°N 16.73°W | 26.92°N 70.90°E |
| 10 | 47.19°N 9.12°E | 47.26°N 11.39°E |
| 11 | 6.16°N 6.79°E | 26.84°N 80.92°E |
| 12 | 52.17°N 5.82°E | 51.93°N 6.42°E |
| 13 | 41.58°S 72.56°W | 41.27°S 173.28°E |
| 14 | 52.19°N 13.86°E | 51.94°N 15.51°E |
| 15 | 40.62°N 44.57°E | 41.11°N 42.70°E |
| 16 | 50.78°N 12.62°E | 48.97°N 14.47°E |
| 17 | 43.98°N 80.06°W | 43.05°N 76.15°W |
| 18 | 49.17°N 8.07°E | 51.51°N 0.13°W |
| 19 | 27.62°N 93.82°E | No result returned |
| 20 | 44.42°N 81.44°W | 43.23°N 86.25°W |
| 21 | 8.07°N 74.80°W | 10.63°N 85.44°W |
| 22 | 43.28°N 80.45°W | 39.95°N 75.16°W |
| 23 | 45.56°N 122.66°W | 41.22°S 174.92°E |
| 24 | 41.52°N 69.96°E | 42.84°N 75.30°E |
| 25 | 30.19°S 70.05°W | 33.37°S 69.15°W |
| 26 | 40.33°N 9.34°E | 41.81°N 6.76°W |
| 27 | 50.32°N 17.60°E | 55.68°N 37.55°E |
| 28 | 42.63°N 8.98°E | 37.38°N 5.97°W |
| 29 | 52.20°N 13.82°E | 54.35°N 18.65°E |
| 30 | 52.23°N 13.93°E | 52.23°N 21.01°E |
| 31 | 51.25°N 5.71°E | 51.44°N 5.48°E |
| 32 | 40.78°N 43.88°E | 42.91°N 23.79°E |
| 33 | 15.64°N 74.11°E | 6.97°N 80.78°E |
| 34 | 52.17°N 13.77°E | 52.23°N 21.01°E |
| 35 | 47.00°N 72.18°W | No result returned |
| 36 | 14.68°N 17.44°W | 34.03°N 5.00°W |
| 37 | 59.31°N 18.03°E | 60.17°N 24.93°E |
| 38 | 33.41°S 70.57°W | 28.27°N 83.97°E |
| 39 | 52.21°N 5.28°E | 52.21°N 5.29°E |
| 40 | 52.15°N 6.00°E | 50.90°N 14.81°E |
| 41 | 52.14°N 106.63°W | 62.24°N 25.72°E |
Why the benchmark changed
The 69-photo set had done its job and had three problems we could not fix by re-running it.
It was too small to settle anything. The 95% interval on 69 photographs was ±8.3 points. Any change we made to the engine disappeared inside it, and a single photograph moved the headline by about 1.4 points. You cannot tell an improvement from noise at that resolution, and we were starting to.
At least one item was unanswerable. One photo in the set — filed under the Spanish city of Granada — is a shot of a ferry deck at sea in fog, with its coordinates recorded well inland in the city itself. Nothing visible in that frame can lead anyone, human or machine, to those coordinates. It was scored as a miss for months. One bad label in 69 is 1.4 points of permanent, meaningless error, and finding it made us audit the rest.
It could only ask one question. The 69 were selected by coordinate as photographs of towns. That made them a fair test of "can you name the town this was taken in" and no test at all of coastline, farmland, forest or road — which is most of what people actually photograph outdoors. The flattering 1.8 km median came directly from that narrowness, and we did not understand quite how much until we built a set that contained something else.
What a photo has to satisfy to be in the new set
- Captured and first published on or after March 1, 2026. Both dates, not just one.
- Camera coordinates only, from EXIF GPS, accurate to within 100 m. Coordinates describing the subject of a photo rather than the camera are excluded, because they measure a different thing and are frequently wrong by kilometres.
- Freely licensed, and capped for diversity: at most three photographs per town and five per uploader, with burst sequences collapsed to one frame, so no single photographer or town can swing the result.
- No duplicates inside the set, and none against the retired 69.
- Outdoor scenes only. Curation removed every indoor shot, food photograph, portrait and close-up.
The date rule is the one worth explaining, because it is the rule that makes the rest of the numbers mean anything. A photograph that has been on the public web for years may well have been in a model's training data, caption and coordinates included. Score an engine on those and you cannot tell reading from remembering — the engine may be recognising the specific picture rather than working out where it was taken. Requiring that a photograph was both taken and first published after training finished closes that door. It is also why the benchmark can only ever be recent photography, and why it has to be rebuilt periodically rather than kept.
How we measure, and what these numbers do not cover
Outdoor photographs only. The set contains no indoor shots, no food, no portraits and no close-ups of objects. Those are a real and substantial share of what gets uploaded to the site, and they are much harder — often impossible — to place. Every figure on this page should be read as accuracy on outdoor photographs, never as accuracy on an arbitrary upload. If you hand the engine a photo of a dinner plate, none of this applies.
The headline is weighted to real traffic. The set deliberately oversamples the rest of the world relative to the uploads we actually receive, because a region needs enough photographs before its number means anything, and left to real proportions Africa and Oceania would have had a handful each. Oversampling fixes that but then makes a raw average describe a world that does not match our users. So the headline reweights each region back to its real share of uploads. It is worth saying that this barely matters here: weighted 91.2% against raw pooled 91.4%. We report both so that no one has to take the weighting on trust.
Africa and Oceania are not reported separately, and will not be. They hold 26 and 27 photographs; South America holds 48. At those sizes the 95% interval is wider than most of the differences anyone would want to read into it, so all three appear only inside the combined rest-of-world row (101 photographs, 91.1%). A per-continent number we cannot stand behind is worse than no number.
Per country, only five reach the reporting floor of 20 photographs. Everything below that floor stays unpublished, including countries where the engine did well.
| Country | Photos | Country accuracy |
|---|---|---|
| United States | 84 | 98.8% |
| Canada | 44 | 86.4% |
| Germany | 32 | 81.3% |
| Spain | 24 | 95.8% |
| Brazil | 22 | 100.0% |
Ten countries in the sampling plan were never sampled at all — no eligible photograph met the rules in time. They are recorded as not sampled. They are not 0%, and if you see them scored as 0% anywhere, that is a mistake and we would like to know about it.
Wikimedia photography is not user photography. Requiring EXIF GPS biases the set toward smartphone and enthusiast photographers, and toward the kind of subject a Commons contributor bothers to upload. It is the best independently verifiable source we have found, and it is still not a sample of our inbox.
The mechanics. Each photograph goes through the deployed engine exactly as an upload on the site does, and countries are scored by ISO country code rather than by matching names. Two analyses errored and are counted as misses. The whole run cost $1.90 — $0.0040 per photograph — and took 23 minutes. This is not a Street View benchmark, so we do not compare it against Street View results such as PIGEON; that is a different task on different imagery. And this is a single sample of 478 photographs: the intervals printed next to every figure are the honest expression of what that many photos can say.
Speed & cost
Three stages, one bill. The median analysis takes 8.5 seconds; nine in ten finish within 11.4 seconds, 95 in 100 within 14.4, and the slowest of the 478 took 26.4 seconds. The figure this post used to publish was about 25 seconds — so the slowest photograph in the current run is roughly what a typical one used to cost you in waiting. Each analysis costs us $0.0040 in model spend, against the roughly one US cent we published before. The demo and accounts are free.
Time per analysis
8.5 s
median · 11.4 s at the 90th percentile, 14.4 s at the 95th, 26.4 s slowest · all three stages
Our model cost per analysis
$0.0040
$1.90 for the whole 478-photo run · what it costs us; the service is free
Response time
Median 8.5 seconds, against about 25 before
Half of the run finished inside 8.5 s and nine in ten inside 11.4 s.
Hover a column, or focus the chart and use the arrow keys, for its exact count.
| Seconds | Photographs, this run |
|---|---|
| 3 to 4 | 5 |
| 4 to 5 | 19 |
| 5 to 6 | 29 |
| 6 to 7 | 45 |
| 7 to 8 | 87 |
| 8 to 9 | 117 |
| 9 to 10 | 79 |
| 10 to 11 | 45 |
| 11 to 12 | 16 |
| 12 to 13 | 5 |
| 13 to 14 | 4 |
| 14 to 15 | 10 |
| 15 to 16 | 7 |
| 16 to 17 | 2 |
| 17 to 18 | 3 |
| 18 to 19 | 1 |
| 19 to 20 | 1 |
| 20 to 21 | 0 |
| 21 to 22 | 0 |
| 22 to 23 | 0 |
| 23 to 24 | 0 |
| 24 to 25 | 1 |
| 25 to 26 | 1 |
| 26 to 27 | 1 |
| Total | 478 |
| This run, median and 90th percentile | 8.5 s and 11.4 s |
| Previously published figure | about 25 s, no per-photograph distribution retained |
Conclusion
v4.0 is live as the engine behind every GeoLocator analysis. It sees with a frontier vision model, decides with a knowledge base built from thousands of real-world clues, and resolves to coordinates with a 34,000-city gazetteer. On 478 recent outdoor photographs it names the right country 91.2% of the time, and it tells you how sure it is in a way that tracks whether it is actually right.
The more useful takeaway is the one that does not fit in a headline: what the engine can do depends far more on your photograph than on the engine. Point it at a street with a sign in it and expect a good answer within a few kilometres. Point it at a forest and expect a country, a region, and a pin that could be a hundred kilometres out. Both of those are in the numbers above, and the reason this post is longer than it was is that averaging them was hiding the more interesting half. Bring a photo of somewhere ordinary and try the demo.
New to this? Start with what geo-estimation is or our guide to finding where a photo was taken.