Zum Hauptinhalt springen

Clinical evidence

This section compiles the evidence behind the automated alopecia scoring provided by Legit.Health: the validation of the model that measures hair loss, the image quality gate, the conference presentations of the method, and the foundational references for the SALT score. Together they document the scientific basis of the automated SALT and what remains to be measured before it is used as a clinical trial endpoint.

Validation status and reproducibility​

A peer-reviewed publication of the automated SALT is in preparation. Until it is published, the evidence is the validation of the hair loss segmentation model under the Quality Management System, reported below, and the method has been presented at national and international dermatology congresses.

Severity scoring has no objective gold standard, and SALT is no exception: it depends on how each rater estimates, by eye, the share of each quadrant that has lost its hair. The model is therefore measured against a pixel-level reference: the consensus of several annotators who traced the boundaries of hair loss on each photograph. Matching that consensus is the realistic performance ceiling for any rater, human or AI.

Two properties are decisive for endpoint use in a trial:

  • Reproducibility: scoring is deterministic. The same image yields the same hair loss percentage at every site and every visit, with no calibration drift and no inter-reader variability in the automated read.
  • Objective extent measurement: SALT is entirely an extent score, and the platform measures that extent by pixel-level segmentation rather than estimating it by eye, so the most operator-dependent step of manual SALT becomes a measurement.

The device is CE-marked as a medical device, meaning a Notified Body has independently assessed it against the safety and performance requirements of the EU Medical Device Directive (MDD 93/42/EEC) and authorised its use on the EU market. Beyond the EU, the device is registered with the MHRA for the United Kingdom market and has obtained ANVISA approval in Brazil. Real-world performance is monitored continuously through the manufacturer's post-market surveillance and post-market clinical follow-up (PMCF) programme under MDR Annex XIV.

Intended use

Legit.Health is a clinical decision support device: the automated scores provide diagnostic support and do not replace the healthcare professional's assessment.

Automated SALT​

The automated SALT combines two steps. A deep learning model segments each of the four scalp photographs into hair, no-hair and non-scalp regions and computes the percentage of hair loss within the visible scalp. The four percentages are then combined with the standard SALT weights (top 40%, back 24%, each side 18%) into the total score. The second step is a fixed formula, so the validation concentrates on the first.

Production model performance​

Each model is validated against its reference under the Quality Management System (IEC 62304 and ISO 14971). The hair loss model was evaluated on an independent test set of 800 images held out from training, covering various types of alopecia, including alopecia areata and androgenetic alopecia. The follicle model is listed for completeness: it works on trichoscopy images and is not part of the SALT calculation.

ModelMetricResult (95% CI)Acceptance criterion
Hair Loss Surface Quantification (EfficientNet-B2 + DeepLabV3+)RMAE7.08% (5.63% to 8.93%)RMAE ≤ 9.6%
Hair Follicle Quantification, trichoscopy (YOLOv11-L)mAP@500.8162 (0.7503 to 0.8686)mAP@50 ≥ 0.72

RMAE (relative mean absolute error) measures, for each test image, how far the model's hair loss percentage falls from the percentage computed from the annotators' consensus mask, relative to that reference, so a lower value is better. It is an image-level error: it is not a difference in SALT points between the AI and a dermatologist. mAP@50 (mean average precision at 50% overlap) measures how accurately the follicle model locates individual follicles, so a higher value is better.

The acceptance criterion is not a target set in isolation. It was derived from the variability between annotators and from the published literature on deep learning measurement of hair loss, including the frameworks of Lee et al. (2020) and Gudobba et al. (2023) in JAMA Dermatology, which showed that the extent of hair loss can be measured automatically from scalp photographs.

Performance across skin phototypes​

Contrast between hair and scalp is the main visual cue for the model, so its error was analysed separately for each group of Fitzpatrick skin phototypes. The mean error stays below the acceptance criterion in every group, and it varies little from the lightest to the darkest skin.

Skin phototypeMetricResult (95% CI)
Fitzpatrick I and IIRMAE6.90% (4.85% to 9.66%)
Fitzpatrick III and IVRMAE7.23% (4.97% to 10.4%)
Fitzpatrick V and VIRMAE7.46% (3.64% to 12.4%)

Each subgroup is smaller than the full test set, so its confidence interval is wider and reaches above the criterion. The consistency of the mean error across groups is the relevant signal.

Robustness to capture conditions​

The model was also tested under image transformations that simulate real-world variation without changing the clinical appearance of the scalp: rotation, changes in brightness and contrast, zoom and reduced image quality. Its performance remained consistent across them, which matters in a multi-site trial where sites use different phones, rooms and lighting.

Fit-for-purpose validation for clinical investigations​

For its use in a Phase 3 programme, the automated SALT workflow was validated as fit for purpose for a clinical investigation, following the FDA guidance on digital health technologies for remote data acquisition in clinical investigations. The validation benchmarked the segmentation against the agreement between annotators on the same images, and documented the design, the data and the performance of the technology for the sponsor.

Image quality for alopecia scoring (DIQA)​

Dermatology Image Quality Assessment (DIQA): Artificial intelligence to ensure the clinical utility of images for remote consultations and clinical trials. Hernández Montilla I, Mac Carthy T, Aguilar A, et al. Journal of the American Academy of Dermatology. 2023;88(4):927–928. DOI: 10.1016/j.jaad.2022.11.002 | PMID: 36526082

Reliable SALT scoring depends on the quality of the input photographs. DIQA is the image quality assessment algorithm that acts as a quality gate in the alopecia imaging workflow, checking every image against consistent quality criteria across investigator sites before it reaches the segmentation model. The dependency is sharp in alopecia, where each visit needs four separate views of the scalp and uneven lighting or shadows can make hair-bearing scalp look like hair loss.

Image quality is the hidden failure point of multi-site imaging. DIQA screens every photograph for clinical utility before it reaches the scoring algorithms, so unusable images are caught at capture rather than surfacing as missing data at database lock.

Dermatology Image Quality Assessment (DIQA): Artificial intelligence to ensure the clinical utility of images for remote consultations and clinical trials. Hernández Montilla I, Mac Carthy T, Aguilar A, et al. Journal of the American Academy of Dermatology. 2023;88(4):927–928. DOI: 10.1016/j.jaad.2022.11.002 | PMID: 36526082

Conference presentations​

AEDV (Spanish Academy of Dermatology and Venereology) 2024, Madrid

Oral: Automated SALT: reducing variability in measuring alopecia areata severity with an artificial intelligence algorithm

Oral presentation of the algorithm that computes SALT automatically from four scalp photographs, aimed at reducing rater variability in measuring alopecia areata severity.

AEDV (Spanish Academy of Dermatology and Venereology) 2025, Valencia

Oral: Automated evaluation of female androgenetic alopecia using artificial intelligence

Oral presentation on the automated assessment of female androgenetic alopecia on the Ludwig scale, a different scale from SALT that uses the same hair loss segmentation.

EADV (European Academy of Dermatology and Venereology) 2025, Paris

Poster: ALUDWIG: An AI-based automated assessment of female androgenic alopecia

Poster on the automated Ludwig grading of female androgenetic alopecia from scalp photographs.

AEDV (Spanish Academy of Dermatology and Venereology) 2026, Maspalomas, Gran Canaria

Oral: Automated quantification of hair follicles using an AI-based object detection model

Oral presentation of the trichoscopy model that locates and counts individual hair follicles and computes follicular density. This model is separate from the SALT calculation.

Foundational SALT references​

The platform does not introduce a new, unvalidated scale. It automates SALT, the measure of scalp hair loss that alopecia areata trials already use, so the endpoint your protocol specifies is the endpoint the algorithm reports.

Olsen EA, Hordinsky MK, Price VH, et al. “Alopecia areata investigational assessment guidelines, Part II. National Alopecia Areata Foundation” J Am Acad Dermatol. 2004. doi:10.1016/j.jaad.2003.09.032

Defines the Severity of Alopecia Tool (SALT): the division of the scalp into four quadrants, their weights (top 40%, back 24%, each side 18%) and the 0 to 100 total score used in alopecia areata trials.

Wyrwich KW, Kitchen H, Knight S, et al. “The Alopecia Areata Investigator Global Assessment scale: a measure for evaluating clinically meaningful success in clinical trials” Br J Dermatol. 2020. doi:10.1111/bjd.18883 PMC7586961

Defines the AA-IGA, five severity categories of the SALT score (none 0%, limited 1 to 20%, moderate 21 to 49%, severe 50 to 94%, very severe 95 to 100%), and reports that clinicians and patients regard reaching 20% scalp hair loss or less as treatment success.

Lee S, Lee JW, Choe SJ, et al. “Clinically applicable deep learning framework for measurement of the extent of hair loss in patients with alopecia areata” JAMA Dermatol. 2020. doi:10.1001/jamadermatol.2020.2188 PMC7489853

An independent group shows that deep learning can determine the SALT score from scalp photographs. One of the two publications used to set the acceptance criterion of the hair loss model.

Gudobba C, Mane T, Bayramova A, et al. “Automating hair loss labels for universally scoring alopecia from images: rethinking alopecia scores” JAMA Dermatol. 2023. doi:10.1001/jamadermatol.2022.5415 PMC9857252

A multicentre study showing that automated percentage hair loss from images can be applied across alopecia types, with errors comparable to annotators. The second publication used to set the acceptance criterion.

King BA, Senna MM, Ohyama M, et al. “Defining Severity in Alopecia Areata: Current Perspectives and a Multidimensional Framework” Dermatol Ther (Heidelb). 2022. doi:10.1007/s13555-022-00711-3 PMC9021348

Reviews how alopecia areata severity has been classified, from the original SALT grades S0 to S5 to the AA-IGA categories, and why there is no single agreed classification.

For the full list of clinical evidence across all indications, see the clinical validation section.

Beschleunigen Sie Ihre Studien

Schließen Sie sich Pharmaunternehmen, Biotechs und CROs an, die 4-6 Monate schneller auf den Markt kommen mit überlegener Datenqualität. Füllen Sie das Formular aus, um Ihre Studienanforderungen zu besprechen.