Ir para o conteúdo principal

Known Limitations

Clinical trial sponsors need to understand not just what an AI scoring system can do, but where its boundaries are. This page documents the known limitations of the automated SALT alopecia severity scoring technology, explains why each limitation exists, and describes how it is managed. Transparency about limitations is a regulatory expectation (ISO 14971, EU AI Act Article 13) and a prerequisite for informed protocol design.

Shadow and parting misclassification​

Natural scalp partings, shadows from overhead lighting, and thin or light-coloured hair can be misclassified as areas of hair loss. The segmentation model uses pixel-level analysis of hair-bearing vs. non-hair-bearing scalp, and shadows and partings share visual features with actual alopecic patches.

Impact: This can lead to overestimation of hair loss, particularly in patients with fine or sparse hair who do not have alopecia, or in images captured under uneven lighting.

Mitigation:

  • The investigator reviews every report and can recapture images when the result does not match the clinical picture
  • The imaging protocol requires consistent, even lighting and specifies that hair should be in a natural position without styling to cover loss
  • The DIQA quality gate rejects images with severe lighting issues before they reach the segmentation model

Clinical agreement has not yet been measured​

Agreement between AI-computed SALT scores and investigator SALT assessments (ICC, MAE, Pearson r) on the same patients at the same visit has not yet been measured. The system is validated at the model level: the hair loss surface quantification model met its acceptance criterion on an independent test set, measured against the annotators' consensus masks (see Clinical evidence). That is an image-level error against a pixel-level reference, not a difference in SALT points against a dermatologist's assessment.

The only prospective clinical investigation of the platform's hair loss scoring to date evaluated female androgenetic alopecia on the Ludwig scale, a different scale from SALT, and did not meet its prospective agreement criterion. It is neither evidence for nor against automated SALT, but it means that clinical validation of automated alopecia scoring is less mature than for psoriasis, atopic dermatitis or hidradenitis suppurativa.

Mitigation: The SALT aggregation is a deterministic weighted sum that adds no model error of its own, so the end-to-end accuracy depends on the segmentation, which is validated. A confirmatory study is committed within the post-market clinical follow-up programme, and agreement with investigator scoring can be built into a study as a sub-study during protocol design. Sponsors should weigh the current evidence when choosing whether automated SALT supports a primary or a secondary endpoint.

The model measures hair loss, not its cause​

The segmentation model classifies each pixel of the scalp as hair, no hair or non-scalp. It measures the extent of visible hair loss; it does not identify what caused it. SALT was designed for alopecia areata, and the model's accuracy has been evaluated on images covering various types of alopecia, including alopecia areata and androgenetic alopecia.

The model has not been evaluated against investigator scores in other types of hair loss, including:

  • Cicatricial (scarring) alopecia: where the scalp texture is permanently altered
  • Telogen effluvium: diffuse thinning without discrete patches
  • Traction alopecia: hair loss from sustained tension on hair follicles

Mitigation: Protocol design should specify which types of hair loss are in scope. Diagnosis remains the investigator's responsibility, and the model's applicability to the study population should be confirmed during protocol design.

Hair texture, colour, and density variation​

Performance may vary with:

  • Very light-coloured hair (blonde, white, grey) on pale scalp: low contrast between hair and skin
  • Very dark dense hair on dark skin: low contrast in the opposite direction
  • Tightly coiled hair textures: different visual presentation of hair density and scalp visibility

The contrast between hair and scalp is the primary visual cue for the segmentation model. Low-contrast scenarios are inherently harder for any image-based assessment method, including human visual estimation.

Mitigation: The hair loss model's error was analysed by skin phototype and met its acceptance criterion in every group, although the darkest phototypes (V and VI) are represented by few images. No analysis by hair colour or hair texture has been reported. The trichoscopy dataset for hair follicle detection contains only Fitzpatrick I and II images, so follicle detection has not been evaluated on darker skin. If the study population includes these groups, this should be discussed during protocol design.

Non-scalp hair excluded​

SALT measures scalp hair loss only. Eyebrow hair loss (assessed by ClinRO Measure for Eyebrow Hair Loss), eyelash loss, and body hair loss are not evaluated by SALT and require separate assessment instruments.

Mitigation: If your protocol requires non-scalp hair assessment, this must be handled separately, either manually by the investigator or with dedicated models if available. This should be discussed during protocol design.

Photograph-based assessment​

The AI analyses photographs of the scalp, not the patient. What only an in-person examination shows, such as a hair pull test, hair fragility or inflammation of the scalp, cannot be read from an image, and hair loss hidden by the camera angle is not measured.

Mitigation: The imaging protocol standardises the four views, the distance and the lighting, and the DIQA quality gate rejects images that do not meet minimum quality standards for focus, lighting, framing and resolution. The investigator reviews every report alongside the clinical examination.

Reference standard​

The reference the model is measured against is the consensus of several annotators who traced the boundaries of hair loss on each photograph. It is the best available approximation of the true extent, but it is not an objective measurement, and the model cannot be more accurate than the reference it was trained and tested against.

Mitigation: Using the consensus of several annotators reduces the bias of any single one, and the acceptance criterion of the model was set from the variability between annotators and from the published literature on automated hair loss measurement.

Model version specificity​

All performance metrics reported in this documentation apply to a specific validated model version. Model updates, including retraining, architecture changes or threshold adjustments, require full re-validation under IEC 62304 before deployment.

Mitigation: The model version is locked at study initiation. No mid-study model updates occur, so every patient in a trial is scored by the same model and endpoint integrity is preserved throughout the study.

Decision support, not autonomous diagnosis​

The system provides severity scoring to support clinical decisions. It does not replace clinical judgement, and it does not make autonomous diagnostic or treatment decisions. All AI-generated scores should be interpreted by qualified healthcare professionals within the context of the patient's overall clinical presentation.

How limitations are managed​

All limitations documented on this page are tracked within the formal risk management process (ISO 14971) and the software development lifecycle (IEC 62304). Each limitation has been assessed for clinical risk, and mitigations have been implemented where the residual risk is not already acceptable.

The post-market clinical follow-up (PMCF) programme under MDR Annex XIV continuously monitors real-world performance. Any new limitation identified through post-market surveillance triggers a formal risk assessment and, if necessary, corrective action.