Research base
Every peer-reviewed study RealSmile cites, mapped to the page that uses it. If a claim on RealSmile is not supported by a study in this list, treat it as opinion, not evidence.
Maintained by Randy, founder of RealSmile. Last verified 2026-05-23.
How to read this page
- Each row is one published study. Authors, year, and venue come straight from the paper.
- The Applied on column lists the RealSmile pages that use that study in body copy. Click through to see the citation in context.
- Reference ranges (averages, percentile distributions) live at /research. The long-form bibliographic write-up of 12 selected studies is at /research/citations.
First Impressions: Making Up Your Mind After a 100-Ms Exposure to a Face
Willis & Todorov (2006) · Psychological Science
Judgments of attractiveness, likeability, trustworthiness, competence, and aggressiveness made after a 100-millisecond exposure to a face correlated highly with judgments made without time limits. From 100 to 500 milliseconds the judgments became more negative and confidence rose; past 500 milliseconds more time mostly raised confidence, and impressions grew more differentiated.
Misleading First Impressions
Todorov & Porter (2014) · Psychological Science
Different photos of the same person produce different first impressions: for a range of social judgments, the variation between images of one person was comparable to or larger than the variation between people. Image-driven preferences appeared after 40-millisecond presentations.
Evaluating Faces on Trustworthiness After Minimal Time Exposure
Todorov, Pakrashi & Oosterhof (2009) · Social Cognition
Trustworthiness judgments made after a 33-millisecond exposure to a face agree with unhurried judgments above chance. Agreement improves with longer exposure and stops improving after about 167 milliseconds.
Applied on
Symmetry, Sexual Dimorphism in Facial Proportions and Male Facial Attractiveness
Penton-Voak et al. (2001) · Proceedings of the Royal Society B
Symmetric male faces were rated more attractive but were not reliably more masculine, and they carry attractive characteristics independent of symmetry that the study could not yet define.
The Evolutionary Psychology of Facial Beauty
Rhodes (2006) · Annual Review of Psychology
A critical review and meta-analyses find averageness, symmetry, and sexual dimorphism attractive in both male and female faces and across cultures.
Applied on
Human (Homo sapiens) Facial Attractiveness and Sexual Selection: The Role of Symmetry and Averageness
Grammer & Thornhill (1994) · Journal of Comparative Psychology
The first study to show that facial symmetry has a positive influence on facial attractiveness ratings, using computer images of faces rated by opposite-sex participants.
Applied on
Facial Action Coding System (FACS)
Ekman & Friesen (1978) · Consulting Psychologists Press
A Duchenne (genuine) smile combines Action Unit 6, orbicularis oculi engagement around the eye, with Action Unit 12, zygomatic major lip-corner pull. Posed smiles use only AU12.
Applied on
Beauty and the Labor Market
Hamermesh & Biddle (1994) · American Economic Review
Holding demographic and labour-market characteristics constant, plain-looking workers earn less than average-looking workers, who earn less than good-looking ones. The penalty for plainness, 5 to 10 percent, is slightly larger than the premium for beauty.
Applied on
Attractive Faces Are Only Average
Langlois & Roggman (1990) · Psychological Science
Composite (averaged) faces are consistently rated more attractive than the individual faces from which they were built. Averageness, not rarity, drives attractiveness.
In Your Face: Facial Metrics Predict Aggressive Behaviour in the Laboratory and in Varsity and Professional Hockey Players
Carre & McCormick (2008) · Proceedings of the Royal Society B
In men, a wider facial width-to-height ratio predicted reactive aggression in a laboratory task and penalty minutes per game in varsity and professional ice hockey players.
Applied on
The Look of a Winner
Todorov & Olivola (2008) · Scientific American Mind
Snap judgments of competence from a single face photo predict real political election outcomes above chance. The effect generalizes to hiring and partner-selection contexts.
Applied on
Evidence from Meta-Analyses of the Facial Width-to-Height Ratio as an Evolved Cue of Threat
Geniole et al. (2015) · PLOS ONE
Across studies, observers judge faces with larger FWHRs as more threatening (r = .46) and more dominant (r = .20), and as less attractive (r = -.26), especially when women make the judgments. The authors call this some support for FWHR as an evolved cue of threat and dominance in men.
Applied on
The Face, Beauty, and Symmetry: Perceiving Asymmetry in Beautiful Faces
Zaidel & Cohen (2005) · International Journal of Neuroscience
Observers shown left-left and right-right mirror composites of fashion-model faces detected asymmetry in them, suggesting that very beautiful faces can be functionally asymmetrical.
Applied on
Anthropometry of the Head and Face
Farkas (1994) · Raven Press, New York
The standard reference for craniofacial anthropometry, defining the head and face landmarks and the measurements taken between them. RealSmile cites it for how a measurement is defined, not for a population average.
SCUT-FBP5500: A Diverse Benchmark Dataset for Multi-Paradigm Facial Beauty Prediction
Liang, Lin, Jin, Xie & Li (2018) · ICPR (arXiv:1801.06345)
5,500 face photos each rated for attractiveness by 60 human raters — the standard academic benchmark for facial beauty prediction. We use it to validate our own scoring: our Impression Percentile model reaches cross-validated r ≈ 0.80 against the averaged human ratings (see "Our own validation" below).
Our own validation · published 2026-07-02
We tested our scoring against 1,566 human-rated faces. Here is exactly what we found.
Method. We ran our production pipeline — the same 68-point landmark detector every scan uses — against the SCUT-FBP5500 academic benchmark (Liang et al., 2018): 5,500 face photos, each independently rated for attractiveness by 60 human raters. We compared our scores to the averaged human ratings using 5-fold cross-validation, so every reported number is measured on faces the model never saw during fitting.
Finding 1 — our original geometry composite failed, and we say so. The 17-metric weighted average we previously used as the overall score showed no positive correlation with human attractiveness ratings (r ≈ −0.33). The individual measurements are real geometry; the blended composite simply does not predict how people rate a face. That is why the report now presents the 17 metrics as measurements — not as an attractiveness verdict.
Finding 2 — the replacement passes. Our Impression Percentile model, built on a 128-dimension facial appearance representation, reaches a cross-validated correlation of r ≈ 0.80 with the averaged 60-rater human judgments (n = 1,566 benchmark faces; consistent across male and female subsets). For reference, published deep-learning models on this benchmark reach r ≈ 0.85–0.90, and simple geometric-feature models reach r ≈ 0.55–0.65.
July 2026 update — v2. A second-generation kernel model, validated the same way, reaches a 5-fold cross-validated correlation of r ≈ 0.836 (percentiles calibrated on out-of-fold predictions). v2 is served to Pro subscribers; v1 (r ≈ 0.80) remains the model in the one-time report.
Scope and limits. The validation covers the Impression Percentile only. It measures agreement with how people rate photos on this benchmark — not dating outcomes, not real-world results, and the benchmark skews toward controlled, front-facing photos. Your number moves with lighting, angle, and expression, which is exactly why the report treats it as photo feedback rather than a verdict on your face.
To our knowledge, no other consumer face-rating tool publishes any validation of its scoring against human ratings. If a competitor publishes theirs, we will link it here.
Read the full Face Score validation write-up — method, per-subset results, and where the number stops applying.
Original data · Photo Variance · published 2026-07-14 · re-measured 2026-08-23
Same face, different photo: how much does the photo move the score?
Among people who scored two or more photos of themselves, the median gap between their highest and lowest Impression Percentile was 13 points, and 513 of the 1,350 — 38% — saw a gap of 20 points or more. Same face, different photo. For a third of repeat scanners, photo choice moved the number more than most facial differences between two strangers would.
What we measured. Every RealSmile scan computes a validated Impression Percentile (0 to 99), a model of how a panel of human raters would place the photo, cross-validated at r ≈ 0.80 against averaged 60-rater judgments (see the validation section above). Since 2026-07-08 the percentile is saved with each scan, so for anyone who scanned more than one photo we can measure how far the same person's number moves between photos. Spread = highest minus lowest percentile across that person's distinct scored photos. Repeat saves of the identical photo are deduplicated first, so the spread compares genuinely different photos.
| Statistic (window: Jul 8 to Aug 24, 2026) | Value |
|---|---|
| Scans with a saved percentile | 6,604 |
| People with 2+ scored photos | 1,350 (4,419 consecutive photo pairs) |
| Median within-person spread | 13 percentile points |
| Spread of 10+ points | 779 of 1,350 people (58%) |
| Spread of 20+ points | 513 of 1,350 people (38%) |
| 75th-percentile spread | 30 points |
| Consecutive photo pairs that moved 10+ points | 1,571 of 4,419 (36%) |
| Scanners (all-time) who scanned more than once | 1,361 of 2,057 |
“Across seven weeks of saved percentile data covering 1,350 repeat scanners, the median repeat scanner's score moved 13 percentile points between photos of the same face, and 38% saw swings of 20 points or more. The photo you choose is a first-impression variable in its own right.”
Why this matters. It is a production-data echo of the published lab finding above: Todorov & Porter (2014) showed that different photos of the same person produce different first impressions, with the variation between one person's images comparable to or larger than the variation between people. Our users reproduce the direction of that effect on their own photos, with a validated score attached.
Honest limitations. RealSmile users are self-selected (skewed young, male, and appearance-focused), not a population sample. People who scan several photos may deliberately vary lighting and angle, which would widen spreads versus casual retakes. Same-day re-saves keep the best score of the day, which can only shrink, never inflate, the reported spreads. None of this measures attractiveness of faces; it measures how much one person's photo-level score varies. We said we would re-publish these numbers as the sample grew and would not delete the early version, so both are here.
Revision, 2026-08-23. The figures above replace the first-week release, which reported a median within-person spread of 7 points from 126 scans and 27 repeat scanners (July 8 to 14, 2026). Same definitions, roughly fifty times the sample, and the spread came out higher rather than lower: 13 points, with the “a third move 20 or more” finding holding at 38%. The original number was not wrong so much as thin — 27 people is too few to pin a median. We are naming the revision rather than restating the table, because a citable page that quietly changes its numbers is worth less than one that shows its corrections.
Cite this as: RealSmile Photo Variance aggregate (July–August 2026), realsmile.online/research-base#photo-variance. Attribution: “RealSmile production data, n = 6,604 scans / 1,350 repeat scanners, July 8 to August 24, 2026.” Aggregates only; no row-level data is published or shared.
Original data · Score Regression · published 2026-08-26
Which way does a face-rating score move on the next photo?
Among people whose first reading of the day landed in the top quartile, 74% of the readings that changed on their next photo came back DOWN. Among people who started in the bottom quartile, 60% went up. Extreme single-photo readings are the least settled ones, in both directions. A very high score is closer to a lucky shot than to a ceiling, and a very low one is closer to an unlucky shot than to a verdict.
What we measured. Every scan pair here is one person, one calendar day, and one scoring-model version — a percentile from an older model is a different ruler and is never paired against a newer one. Within that, the first reading is compared to the next genuinely different photo: taken more than 30 seconds later, and not an identical reading, which removes same-capture repeats. Held at one pair per distinct person, so someone who scans on six days contributes once rather than six times. Window: June 1 to August 26, 2026.
| First reading | People | Reading changed | Of those, went UP | Typical move |
|---|---|---|---|---|
| Bottom quartile (below 25th) | 143 | 85 | 60% | +4 points |
| Middle half (25th–74th) | 716 | 458 | 46% | −1 point |
| Top quartile (75th and above) | 457 | 281 | 26% | −8 points |
What carries the result. The ordering across the three rows, not any single row. The top-quartile figure is overwhelming — 26% up against a coin flip is p ≈ 1×10-16 on an exact binomial — while the bottom-quartile 60% is suggestive rather than conclusive on its own (p ≈ 0.08, n = 85). Read together they are the signature of regression to the mean, which is what an unstable measurement does at both ends. A floor effect would lift the bottom row alone and leave the top row untouched; a ceiling and a floor are different mechanisms, and one gradient spanning both is not either of them.
Why exact repeats are set aside. The percentile is a whole number, so a substantial share of second photos return the identical value and carry no directional information at all. Those are counted in the “people” column and excluded from the direction split, which is stated rather than buried because including them either way would change the headline: read across ALL bottom-quartile pairs including ties, only 39% rose. “Low scores usually go up” is not true and is not what this says.
“Across 1,316 people who photographed themselves twice in one day, face-rating readings regressed toward the middle: 74% of changed readings that started in the top quartile came back down, against 40% of those that started in the bottom quartile. The more extreme a single-photo score, the less of it survives the next photo.”
Cite this as: RealSmile Score Regression aggregate (June–August 2026), realsmile.online/research-base#score-regression. Attribution: “RealSmile production data, n = 1,316 people, one same-day photo pair each, June 1 to August 26, 2026.” Aggregates only; no row-level data is published or shared.
Why this page exists
The face-perception literature is well-established but spread across psychology, anthropology, and economics journals. A single lookup that maps each study to the RealSmile page that uses it makes the citation chain checkable in one click. The studies above are the load-bearing references for the 17-metric audit, the dating-photo audit, the LinkedIn headshot audit, and the four long-tail measurement guides.
For the original-data companion to this index, see the State of Looksmaxxing 2026 data report, which publishes RealSmile's aggregated face-scan distributions alongside the published-research baselines listed above.