What we measure, and why order comes first

The point of dating scanned photos is a timeline that reads in the right direction: the wedding before the baby photos, the baby photos before the graduation. So our primary metric is pairwise chronological ordering: take every possible pair of photos in a test set, ask whether the AI’s dates put the earlier photo first, and report the percentage of pairs it gets right.

We track three supporting metrics alongside it:

  • Wrong-decade rate. How often a photo’s estimated date lands in a different decade than the true date. Decade-scale misses are what make a timeline feel broken.
  • More-than-10-years-off rate. The catastrophic misses. Our benchmark bar for promoting any change to the dating prompt is that this stays at 0% on the realistic archive sets.
  • Within-1-year rate. How often the estimate lands within a year of the true date.

A single photo’s exact year is genuinely hard, even for a person; relative order is what the evidence supports best, which is why we lead with it and treat exact-date accuracy as secondary.

How the benchmark works

We test against ground-truth sets: real scanned family photo collections (63 to 375 photos in the realistic sets, plus two frozen 200-photo hardest-case sets) where the true date of every photo is known from printed lab timestamps, handwritten notes on the backs, and family records, verified by a person before the set is frozen. The AI is scored against those known dates; it never sees them.

The sets are split by difficulty. The realistic archive sets (prints spanning 1986 to 2005) look the way customer archives actually look: some prints carry lab-printed dates or dated backs, the rest carry nothing, and the system reads whatever evidence each print offers. The bare-front sets are the hardest case: two frozen 200-photo sets (prints spanning 1968 to 2006) selected so that no print carries any printed date, no photo backs exist, and every filename is a meaningless scanner name, so the model must judge each photo purely from visual era clues (film stock, print borders, clothing, cars) and the apparent ages of the people in it.

The underlying vision model is not deterministic, and a single run can swing a few points. Every number we publish comes from repeated runs of the same frozen benchmark (we report ranges across runs, never a single best run), and we re-run the benchmark before any change to the production dating prompt ships. A change is promoted only if ordering holds, the wrong-decade rate does not regress, and the more-than-10-years-off rate stays at 0% on the realistic sets.

The three evidence levels, at a glance

Most evidence

Realistic archives

Some prints carry a lab or camera date, some backs carry a handwritten note, and album folders have real names. The system reads all of it.

  • 1994 Emma’s birthday
    • scan_012.jpg
    • scan_013.jpg
  • Disneyland trip 1998
Order only

Bare fronts, album order

No printed dates, no backs, meaningless scanner names. But the albums are uploaded in order, so neighboring photos and how people age still carry clues.

  • Album 03order kept
    • Scan_0141.jpg
    • Scan_0142.jpg
    • Scan_0143.jpg
No evidence

Bare fronts, shuffled

The same bare prints as loose singles in random order. The model is left with visual era clues alone: film stock, print borders, clothing, cars.

  • Scan_0047.jpg
  • IMG_3391.jpg
  • Untitled_12.jpgno order, no dates
no tags people tagged with birth years, scored on the photos that contain them All bars share a 0–100% scale except the last row, drawn on a 0–3 year scale.

Any two photos in the correct chronological order

Realistic archives
97.7–99.2%
no change
Bare fronts, album order
About 65%
no material change
Bare fronts, shuffled
75–80%
84%

Our primary metric: of the photo pairs the model dates apart, the share where the earlier photo comes first. On album-order uploads the model often gives a whole album one shared estimate, so fewer pairs get dated apart there and the share swings widely from run to run (46–79% across runs). Tagging does not reliably move ordering on the top two tiers. On shuffled prints it helps the photos it can see into: those same photos go from 78% to 84%, in all three runs.

Estimated within 1 year of the true date

Realistic archives
96–100%
no change
Bare fronts, album order
72%
no change
Bare fronts, shuffled
~25%
51%

Tagging moves this only where nothing else carries it: on shuffled prints, the photos with a tagged person in them roughly double, 23% to 51%, in every run. Ordered albums already carry within-a-year on their own, so tags neither add nor cost anything there (72% to 74%, inside the run-to-run swing). Counted across all 200 shuffled prints, two thirds of which contain nobody tagged, those same runs read about 25% to about 38%.

Exact year correct

Realistic archives
81–96% per set
no change
Bare fronts, album order
33.5%
55%
Bare fronts, shuffled
10–17%
31%

The album-order set is where tags pay off: the photos with a tagged person go from 35% to 55%. On the shuffled set the same photos go from 18% to 31%. The album-order gain is the one that travels: because a pinned person anchors the photos scanned around them, prints in those albums with nobody tagged in them move almost as far (34% to 53%), which does not happen on shuffled singles.

Placed in the wrong decade (lower is better)

Realistic archives
under 1%
no change
Bare fronts, album order
2.5%
no change
Bare fronts, shuffled
30–45%
23%

Decade placement is driven by visual era clues (film stock, print borders, clothing, cars), so tags do not reliably move it on the top two tiers. On shuffled prints they do reach the photos they can see into: a birth year rules out whole decades at once, and those photos fall from about 35% misplaced to about 23%. Every run improved, though by anywhere from 3 to 20 points.

Placed more than 10 years off (lower is better)

Realistic archives
0.0% in every run
still 0
Bare fronts, album order
0.0%
still 0
Bare fronts, shuffled
0.5–4.5%
no change

The catastrophic misses, and our promotion bar: no change to the dating prompt ships unless this stays at 0% on the realistic sets.

Median distance from the true date (lower is better; 0–3 year scale)

Realistic archives
days, not months
no change
Bare fronts, album order
about 1 year
about 3 months
Bare fronts, shuffled
about 2–3 years
~1 year

Realistic archives sit at days rather than months (0–13 days median across runs) because the printed timestamps give exact days. Counting only the photos with a tagged person in them, the album-order set’s typical miss falls from about a year (363 days) to about three months (92 days), and the shuffled set’s from about two and a half years (864 days) to about one (365 days).

Sets: realistic archive sets (63–375 photos per set, prints spanning 1986–2005, printed timestamps or dated backs on part of each set) and the two frozen 200-photo bare-front sets (1968–2006). The striped “people tagged” bars are the July 2026 face-tagging benchmark: each tier’s set run three times with and without six family members tagged with birth years. Those bars count only the photos that contain one of the tagged people (70 of the 200 prints in each bare-front set, 179 of the 375 realistic prints), pooled across the three runs, because a photo with nobody tagged in it has nothing to gain and would only dilute the number; the solid bars are the whole set. On the realistic set the tagged and untagged runs are statistically identical on those photos too, which is why every striped bar in that column reads no change; the bare-front measurements are in the next section. Bare-front runs swing far more from run to run than realistic-archive runs; ranges are across repeated runs, never a single best run.

In plain terms: when your prints carry the evidence real archives usually carry (a printed timestamp on some prints, a note on some backs, album folder names), the timeline reads in the right direction and nearly every photo lands within a year of the truth. When a photo offers no evidence at all, the model can still place the era, but exact years are genuinely hard, and upload order matters: albums scanned in order give the model relative-age and event clues that shuffled single prints do not. Run-to-run swings on the bare-front sets are large, so treat those figures as indicative rather than precise. That hardest case is exactly why the review timeline exists, why the system reads every printed date and photo back for you instead of relying on visual guesses, and why we built face tagging with birth years: on shuffled bare prints it roughly doubles the within-a-year rate, and on ordered albums it roughly doubles the exact-year rate, both measured in the next section.

Messy archive with no dates? Tag the people

The shuffled bare prints above are the worst case, but they are not a dead end, because most family archives have one more source of evidence: the family. Timeline Scan finds the faces across your archive and groups them by person; name a person once, add their birth year, and every photo they appear in gains hard evidence: who is in the frame and how old they look. A tagged person born in 1991 who looks about three pins a bare print near 1994, and no photo of them can ever be dated before they were born.

A scanned 1990s family photo of three-year-old Emma with her father Marcus
Emma · born 1991 Marcus · born 1962

No printed date, no back, no filename. But a three-year-old Emma pins this print near 1994.

Land within 1 year of the true date (photos with tagged people, shuffled bare-front set)

No tags
23%
People tagged with birth years
51%

Typical distance from the true date on those photos: about 2.4 years without tags, about 1 year with them. Exact-year rate goes from 18% to 31%.

We measured this in July 2026 on the shuffled bare-front set, the hardest case on this page, by running the same benchmark three times with and without tags (70 of the 200 prints contain people from the tagged family). On the photos with tagged people, the within-a-year rate roughly doubled, from 23% to 51%, and the typical miss fell from about two and a half years to about one. Those same photos also landed in the right decade more often (35% misplaced down to 23%) and in the right order more often (78% to 84% of the pairs the model dated apart). Diluted across all 200 prints, most of which hold no tagged face at all, that same gain reads as about 25% to about 38%, a lift of 9 to 17 points that held in every run. Adding only a family list with birth years, without tagging faces in the photos, delivered about a third of that gain.

We then ran the same three-run benchmark on the album-order set, again counting the photos with tagged people in them. There the within-a-year rate does not move (about 72% either way: the album order already carries it), but tagging still sharpens the result where order alone cannot: the exact-year rate roughly doubles, from 35% to 55%, and the typical miss falls from about a year to about three months. And unlike the shuffled set, the gain spreads across the whole album rather than staying on the photos with tagged people, because one pinned person anchors every photo scanned around them: the prints in those albums with nobody tagged in them went from 34% to 53% on exact year, nearly the same jump.

Two honest qualifiers. On shuffled loose prints the gain concentrates on photos that actually contain the people you tag; prints of scenery, or of people outside the family, do not benefit directly (on ordered albums the anchors reach further, as above). And photos that already carry strong evidence do not need the help: we ran the same three-run benchmark on the realistic 375-photo set, and with or without tags the results are statistically identical, on the 179 photos with tagged people just as much as across the whole set. Tagging is how you rescue the photos that have nothing else going for them.

Honest limitations

  • The ground truth comes from real family archives, not a public dataset. That makes it representative of actual customer photos (mixed eras, mixed print quality, duplicates, backs with and without notes), but it is not a benchmark anyone can independently re-run, and we say so.
  • Our verified ground truth spans 1968 to 2006. The realistic archive sets cover 1986–2005 and the bare-front sets cover 1968–2006. The dating rules cover earlier eras (card-mounted portraits, deckled-edge black and white, early color), but accuracy on photos from before the late 1960s is designed-for rather than measured. We’re building an older ground-truth set to close that gap.
  • Exact years are harder than order. If you need every photo within a year, expect to review and nudge some estimates; the review timeline exists for exactly that.
  • Run-to-run variance is real. Single benchmark runs of the same system vary by 1–3 points on the realistic sets and by far more on the bare-front sets, which is why we publish ranges across repeated runs, not best runs.

For comparison, MyHeritage reports that PhotoDater’s year estimate is correct about 60% of the time on its own test data. Our exact-year rate runs 81–96% on realistic archive sets and 10–33.5% on bare fronts with no evidence at all, which is the honest way to read any photo-dating accuracy number: it depends almost entirely on how much evidence the photos carry. Our side-by-side comparison covers when each tool fits.

Check the accuracy on your own photos

The free photo trial exists so you can judge the dating on your own scans before paying anything.

Try It Free