Our official iOS app is live on the App StoreDownload free

Photo Geolocation API: What It Can and Can’t Tell You

· 11 min read

Short answer

A photo geolocation API accepts an image and returns an estimate of where it was taken: a country, region or city, approximate coordinates and a confidence score. The estimate is inferred from visible evidence such as architecture, signage and vegetation. It is not a GPS reading, and pixels alone cannot produce a street address.

The name causes confusion before any code is written. The browser’s Geolocation API reports where a device is right now, and only after the user grants permission. A photo geolocation API answers a different question about a different input: where was this picture taken? The picture may be ten years old and from anywhere on Earth.

Three neighbouring technologies also get described as "locating a photo": reading GPS metadata, reverse image search and landmark detection. They return different things and fail in different ways, so it pays to know which question your product is really asking before you choose one.

What does a photo geolocation API actually do?

It infers a location from the pixels. A model reads evidence such as road markings, architecture, plants, the script on signs and the quality of the light, then reports the place where most of that evidence agrees, with a score for how strongly it agrees.

Many research systems treat this as a classification problem. PlaNet, a 2016 model from Google researchers, divided the surface of the Earth into thousands of cells and trained a network on millions of geotagged photos to predict which cell an image belongs to. The output is a probability spread across the globe. A photo of a Lisbon tram packs that probability into one place; a photo of a pine forest spreads it thin across several continents. A confidence score is, in effect, a summary of that spread.

The detail that matters most for product design is that nothing has to exist online. The API works on a photo that has never been published, and it can return an answer even when the evidence is weak, which is why the score matters as much as the place. How a model reads a scene is covered separately; this guide is about what you do with the output.

Precision is also uneven by design. PlaNet gave densely photographed areas smaller cells than sparse ones, so it could reach street-level accuracy in some city areas while rural regions sat in much larger cells. Expect most models to be sharper where their training photos were dense.

How is it different from EXIF GPS, reverse search and landmark detection?

EXIF GPS is a measurement stored in the file. Reverse image search finds copies of the image online. Landmark detection recognises famous structures. A photo geolocation API infers a region from the scene itself, the only option for an ordinary street photo with no metadata and no copies online.

MethodWhat it readsWhat it returnsFails when
EXIF GPSCoordinates the device wrote into the fileA measured point, typically within 5 m under open skyThe metadata was stripped, as most platforms do on upload
Reverse image searchCopies of the image already on the webMatching images and the pages they appear onThe scene has never been published
Landmark detectionWell-known natural and human-made structuresLandmark name, score and the landmark’s coordinatesNo famous structure is in frame
Photo geolocationVisual evidence across the whole sceneCity, region or country estimate with a confidence scoreThe frame holds no geographic evidence
Four ways to get a location from an image

EXIF GPS is a measurement, not an estimate

When a phone records location, the coordinates go into a dedicated GPS block of the EXIF data, under tags such as GPSLatitude and GPSLongitude. GPS.gov puts typical smartphone accuracy at 4.9 m under open sky. No visual method comes close, so a sound pipeline reads the metadata first and only infers when it is missing. For images that have passed through a social platform it usually is, because re-encoding on upload drops it.

Reverse image search is a lookup

Google Cloud Vision’s web detection shows the shape of this category. It returns fields such as fullMatchingImages, partialMatchingImages and pagesWithMatchingImages. When a match exists, the page it sits on often names the place outright, which beats any estimate. When the photo is private or original, there is nothing to match. The engine comparison covers the consumer versions of the same idea.

Landmark detection recognises famous places

Cloud Vision’s landmark detection "detects popular natural and human-made structures within an image". Each result carries a description with the landmark’s name, a score, a boundingPoly marking where it sits in the frame, and a locations list of coordinates. Google’s own example returns Saint Basil’s Cathedral with a score of 0.78.

Two details catch teams out. An ordinary street returns no landmark at all, so coverage is narrow. And the coordinates can describe the landmark rather than the camera: Google’s reference notes that one location may describe the scene in the image while another describes where the photo was taken. Someone photographing the Eiffel Tower from across the Seine may stand several hundred metres from the tower’s own coordinates.

Photo geolocation estimates a region for any outdoor scene

Visual inference is the only one of the four with something to say about an unremarkable street that has never been online. The cost is precision. The answer is a city or region with a confidence score, never a measured point, and a frame with no geographic evidence gets a weak answer or none.

What does a photo geolocation API return?

Usually a place estimate at one or more levels, such as country, region and city, plus approximate coordinates that stand for that estimate and a confidence score. A well-designed response also reports GPS metadata separately when the file still carries it, so a measurement is never mistaken for a guess.

FieldWhat it meansHow to use it
Country, region, cityThe place the evidence points to, at several levelsShow the most specific level the confidence supports
Approximate coordinatesA point standing for the estimated area, not the camera positionShow a city label or an area, never a street-level pin
Confidence scoreHow strongly the visual evidence agreesSet thresholds against your own labelled test photos
EXIF GPS, when presentCoordinates measured by the device and stored in the fileKeep it in its own field; it outranks any estimate
Typical response fields and how to handle them

The design mistake to avoid is merging the estimated coordinates and the EXIF GPS into one location field. Downstream code, and the people reading its output, then treat a city-level guess with the confidence of a satellite fix. Keep measured and inferred locations apart all the way to the interface, and label them differently when you show them.

Estimated coordinates need the same care. A point returned for Porto stands in for an area of many square kilometres. Plot it as a pin at street zoom and users will read a precision that was never there.

What will the GeoSpy AI API return?

The GeoSpy AI API is not open yet. It is being designed to return the same fields the website analysis returns today: a city, region and country estimate, a confidence score, approximate coordinates for the estimate, and EXIF GPS coordinates when the uploaded file still contains them.

That list maps directly onto the table above. It contains no street address, because visual evidence cannot support one, and it lists the metadata reading next to the visual estimate rather than in place of it.

Endpoints, pricing, rate limits and a launch date have not been announced, and nothing on this page should be read as a promise about them. If you are planning an integration, request early access through the contact page and describe your use case.

Until then, the website is the closest preview of the output on real photos. Web analysis comes with GeoSpy Premium, a subscription of 7.99 US dollars a week or 29.99 US dollars a year that renews until you cancel, and uploaded photos are not stored.

What is a photo geolocation API used for?

Mostly for consistency checks and triage: flagging a photo whose likely location disagrees with a claim, so that a person can review it. Marketplaces, newsrooms, travel and photo apps, insurers and researchers all have versions of that question. It is about places, never about individuals.

  • Marketplace listing checks. A holiday rental advertised in Barcelona whose photos point strongly to another country deserves a second look. Send the mismatch to a reviewer rather than rejecting the listing automatically, and remember that interior shots, which make up much of any listing, give a model little to read.
  • Newsroom verification. A viral image said to show one city can be tested against the evidence in the frame before anyone publishes it. Geolocation answers only "where", so pair it with a search for earlier copies, as laid out in how to verify a photo.
  • Travel and photo apps. Scanned prints and exported images often carry no GPS. An estimate can suggest an album location or group photos into a trip, as long as the interface presents it as a suggestion the user confirms.
  • Insurance claim review. When a claim places damage in one town and the photos point clearly elsewhere, an adjuster has a reason to ask questions. The estimate alone is not grounds to deny anything.
  • Research. Mapping the geographic spread of a large image archive, or auditing a training set for regional bias, needs coarse labels at scale, which is exactly what the method produces well.

The pattern across all five is the same: the estimate flags, and a person decides. A product that lets a city-level guess trigger an automatic outcome inherits every error the model makes, at scale. A structured OSINT workflow shows where an estimate fits among the other checks.

What can’t visual geolocation do?

It cannot produce a street address from an ordinary scene, place interiors or close-ups reliably, date a photo, identify anyone in it, or confirm that the image is genuine. It can also be confidently wrong when a landscape or building style occurs in many countries.

Some of these limits belong to the evidence, not the model. The PlaNet paper opens by noting that it is trivial to construct situations where no location can be inferred. A white wall, a hotel bathroom or a plate of food carries no geography, and no increase in model size changes that.

Research results are a useful reality check. On 2.3 million Flickr photos, PlaNet placed 10.1% within 25 km, the distance its benchmark counts as city level. PIGEON, presented at CVPR 2024 and trained on Street View imagery from GeoGuessr, placed over 40% of its guesses within 25 km. Part of that gap is newer methods, and part is the photos. Street View frames are full of road markings and signs, while a general photo collection contains many frames with nothing geographic in them.

Geolocation also cannot vouch for an image. An AI-generated street scene still receives an estimate for the place it most resembles. The API answers "what place does this look like", not "is this real".

How should you test an API before integrating it?

Use your own photos with GPS metadata removed, and score results at fixed distances, such as 25 km for city level and 750 km for country level. Include featureless photos on purpose, and check that high-confidence answers really are right more often than low-confidence ones.

  1. Step 1Build a test set from real trafficDemo images are chosen to succeed. Collect a few hundred photos that resemble what your users will send, with known true locations, and keep the set fixed so vendors and versions can be compared fairly. Anecdotes mislead too: a viral 1,100-word prompt for ChatGPT impressed many who tried it, yet on a fixed set of 200 photos it did no better than a short one.
  2. Step 2Strip GPS before you benchmarkIf the API also reads metadata, a geotagged test photo measures the metadata reader rather than the model. Remove EXIF from the test copies, or you will record a perfect score that vanishes in production.
  3. Step 3Score at fixed distancesResearch benchmarks count a guess as correct within 1 km (street), 25 km (city), 200 km (region), 750 km (country) or 2,500 km (continent). An accuracy claim with no distance attached cannot be compared with anything.
  4. Step 4Mix cities and countrysidePrecision tends to follow the density of training photos, so a model that is sharp in capitals can be vague in farmland. Test both if your users send both.
  5. Step 5Include photos with nothing to readInteriors, close-ups and plain skies should come back with low confidence. A tool that names a city from a hotel room with high confidence will do the same to your users.
  6. Step 6Check that the score is calibratedGroup results by confidence and measure how often each group was right. If answers scored above 0.8 are not clearly more accurate than answers near 0.4, the score cannot drive thresholds.
  7. Step 7Read the data termsFind out whether images are stored, for how long, and whether they are used for training. Photos can count as personal data, especially when people appear in them.

Frequently asked questions

Is a photo geolocation API the same as the browser Geolocation API?
No. The browser API reports where the user’s own device is, and only after the user grants permission. A photo geolocation API works on an image, which could have been taken anywhere, at any time, by anyone.
Can a photo geolocation API give an exact address?
Not from visual evidence alone. Two streets a kilometre apart usually look alike, so an honest result is a city or region. Exact coordinates come only from GPS metadata in the file, when it has survived.
Does visual geolocation work on screenshots?
A screenshot carries no GPS from the original photo, as the screenshot guide explains. Visual inference still works if the screenshot shows an outdoor scene at a usable size.
When will the GeoSpy AI API be available?
No date has been announced. Developers and product teams can request early access now. Until it opens, the website analysis returns the same kinds of fields the API is being designed around.
Can a photo geolocation API identify people?
No. It estimates where a scene is, not who appears in it, and it should never be used to locate or follow anyone.

Sources

  1. Detect landmarks — Google CloudDefinition, response fields, and the Saint Basil’s Cathedral example scored 0.78.
  2. AnnotateImageResponse — EntityAnnotation — Google CloudScore range 0–1; a location may describe the scene or the place the photo was taken.
  3. PlaNet — Photo Geolocation with Convolutional Neural Networks — Weyand, Kostrikov and Philbin (arXiv)Benchmark distances of 1, 25, 200, 750 and 2,500 km; 10.1% city-level accuracy on 2.3M Flickr photos.
  4. PIGEON: Predicting Image Geolocations — Haas, Skreta, Alberti and Finn (CVPR 2024, arXiv)Over 40% of guesses within 25 km on Street View imagery.
  5. GPS Accuracy — GPS.govSmartphones are typically accurate to within 4.9 m under open sky.

Keep reading

Çerez bildirimi

Bu site; oturum, dil ve güvenlik için gereken zorunlu çerezlerin yanında, trafiği ve reklam performansını ölçmek için Google Analytics ve Meta Pixel çerezlerini kullanır. Bunları tarayıcı ayarlarınızdan engelleyebilirsiniz; ayrıntılar Çerez Politikası'nda.

Cookie notice

Alongside the strictly necessary cookies for sign-in, language and security, this site uses Google Analytics and Meta Pixel cookies to measure traffic and advertising performance. You can block them in your browser settings; details are in the Cookie Policy.