Clinical Evidence and Validation
ALADIN combines lesion count and density into an IGA-aligned severity score. Its peer-reviewed evaluation puts the results alongside expert dermatologist assessments, giving your team a transparent basis for evaluating image-based acne scoring.
How the ground truth is established
The study evaluates ALADIN and individual dermatologists against the same expert-consensus reference. This gives sponsors a clinically relevant comparison: how closely does automated scoring agree with a pooled expert assessment, and how does that compare with individual readers?
AcneSeverity-V1 supported formula optimisation and in-sample evaluation. AcneSeverity-V2 added an independent, exploratory evaluation on public-atlas images, extending the assessment to a dataset with predominantly severe acne.
Understanding the metrics
Four metrics are used to evaluate the agreement between ALADIN scores and the expert consensus. Together, they show how scores relate to the consensus and the average size of scoring differences. The displayed bands help interpret these metrics; they are descriptive guides:
Pearson correlation
The strength and direction of the linear relationship between two sets of scores. In this context: how closely ALADIN's scores track the expert consensus scores.
Perfect score: 1.0 would mean perfect linear correlation with the consensus. However, individual dermatologists themselves only achieve ~0.56–0.65 against the consensus, because acne grading is inherently subjective.
Why it matters: Shows how closely automated severity scores follow the linear pattern of expert-consensus scores across the evaluated images.
Spearman correlation
The strength of the monotonic (rank-order) relationship between two sets of scores. Less sensitive to outliers than Pearson. Measures whether ALADIN ranks patients in the same severity order as the consensus.
Perfect score: 1.0 would mean perfect rank agreement. Individual dermatologists achieve ~0.58 against the consensus.
Why it matters: Shows how consistently automated scoring orders the evaluated images from lower to higher severity compared with the expert consensus.
Cohen’s kappa (κ)
Agreement between two raters beyond what would be expected by chance. Uses quadratic weights, meaning a 2-grade disagreement is penalised more than a 1-grade disagreement. This is the primary measure of clinical agreement.
Perfect score: 1.0 would mean perfect agreement. The dermatologist results give you a clinical comparator for the automated scores.
Why it matters: Measures agreement with the expert consensus while accounting for chance and giving greater weight to larger scoring disagreements.
Mean absolute error
The average magnitude of scoring errors in IGA points. An MAE of 0.63 means that, on average, ALADIN’s score differs from the consensus by less than 1 IGA grade.
Perfect score: 0.0 would mean zero error. Individual dermatologists achieve MAE of 0.44–0.65 against the consensus — even experts don’t agree perfectly.
Why it matters: Puts scoring differences into familiar IGA units: the average absolute distance from expert consensus across all evaluated images. Read it alongside correlation and kappa to assess both agreement and error.
Severity scoring performance: ALADIN vs. dermatologists
The following results are from the peer-reviewed validation study (Sabater et al., Skin Health and Disease, 2026). They evaluate the ALADIN scoring formula, the mathematical model that converts lesion count and density into an IGA-aligned severity score. Both ALADIN and individual dermatologists are measured against the same ground truth (dermatologist consensus). Values in bold with a check mark indicate a numerically equal or more favourable point estimate: higher for correlation and kappa, lower for MAE. Check marks highlight numerical comparisons, not statistical significance.
IGA agreement
| Metric | Mild-to-moderate | Severe cases | ||
|---|---|---|---|---|
| ALADIN | Dermatologists | ALADIN | Dermatologists | |
| Pearson | 0.58 ✓ | 0.56 | 0.65 ✓ | 0.65 |
| Spearman | 0.56 | 0.58 | 0.54 | 0.58 |
| MAE | 0.63 (0.49) ✓ | 0.65 (0.63) | 0.64 (0.71) | 0.44 (0.5) |
| Cohen’s κ | 0.53 ✓ | 0.46 | 0.61 | 0.62 |
How to read this table
- Agreement with expert assessment: Correlation and weighted kappa let you compare ALADIN with individual dermatologists against the same consensus. The independent dataset shows similar correlation and kappa estimates, alongside a higher device MAE.
- Average scoring difference: MAE expresses the average absolute difference from consensus in IGA grades across all evaluated images. ALADIN has lower MAE on AcneSeverity-V1; dermatologists have lower MAE on AcneSeverity-V2.
- Study context: V1 was used for optimisation, while V2 provides an independent, exploratory comparison. The study evaluates agreement on individual images rather than changes over a course of treatment.
Put the evidence to work for your trial
Review the scoring method, expert comparisons and image acquisition requirements together to assess how automated scoring fits your study. The published results make that assessment transparent: you can examine both agreement and error, with separate results for each dataset.
Explore the trial workflow to see how image capture, scoring and data export fit together, or review the scoring methodology for the calculation behind each endpoint.
Lesion detection performance
The following metrics are validated per the Quality Management System (IEC 62304 / ISO 14971).
- mAP@50 (95% CI): 0.45
- rMAE Papule (95% CI): 0.58
- rMAE Pustule (95% CI): 0.28
- rMAE Comedone (95% CI): 0.62
- rMAE Nodule/Cyst (95% CI): 0.33
How to read these metrics
- mAP@50 = 0.45 (95% CI: 0.43–0.47): The inflammatory lesion detector achieves mean average precision of 0.45 at the standard 50% IoU threshold. The acceptance criterion (≥0.21) is based on non-inferiority to published acne detection studies, where the literature range is 0.21–0.54.
- rMAE (relative Mean Absolute Error): For per-type detection, each lesion type's counting error is compared against the inter-rater variability of expert dermatologists on the same task. The acceptance criterion is that the model's rMAE must not exceed the experts' rMAE. All lesion types pass, meaning the AI counts each lesion type at least as consistently as dermatologists disagree with each other.
The lesion detector is not used in isolation; it feeds into the ALADIN/IGA formula, which was co-optimised with the detector. The formula constants (, ) were calibrated against reference IGA scores using the detector's actual output. The publication reports end-to-end IGA agreement for the detector and formula together. This connects the component-level measurements to the severity score evaluated against expert consensus.
Additionally, the acceptance criteria are based on non-inferiority to expert inter-rater variability, a stronger validation methodology than simple accuracy benchmarks. The question is not "is the AI perfect?" but "is the AI at least as consistent as dermatologists are with each other?"
Validation datasets
| Dataset | Images | Source | Severity profile | Annotators | Role |
|---|---|---|---|---|---|
| Mild-to-moderate | 331 | Private (Clínica Dermatológica Internacional, Spain) | Predominantly mild-to-moderate facial acne (IGA 0: 14, IGA 1: 65, IGA 2: 150, IGA 3: 83, IGA 4: 19) | 3 | Primary dataset for ALADIN formula design and evaluation |
| Severe cases | 27 | Public dermatology atlases (DermAtlas, Atlas Dermatológico) | Predominantly severe acne (IGA 0: 1, IGA 1: 3, IGA 2: 10, IGA 3: 11, IGA 4: 2) | 2 | Exploratory test set for evaluating generalisation to severe cases and third-party images |
The two datasets complement each other:
- 331 images: AcneSeverity-V1, used for formula optimisation and in-sample evaluation, predominantly mild-to-moderate acne, annotated by 3 specialists
- 27 images: AcneSeverity-V2, an independent exploratory public-atlas dataset, predominantly severe acne, annotated by 2 specialists
ALADIN peer-reviewed publication
Sabater A, et al. “ALADIN: Acne Lesion And Density INdex. A Novel Tool for Automatic Acne Severity Assessment” Skin Health and Disease (British Association of Dermatologists). 2026. (Provisionally accepted)
- Development of the IGA = N^a · (D + b) formula combining lesion count and spatial density into an IGA-aligned severity score
- Calibration of formula constants against the consensus of three board-certified dermatologists
- Validation of the CNN-based inflammatory lesion detection model (papules, pustules, nodules)
- Evidence that incorporating spatial density alongside lesion count improves alignment with clinical severity perception
The paper was presented as a poster at the AEDV 2025 conference in Paris prior to journal submission.
DIQA validation
Hernández Montilla I, Mac Carthy T, Aguilar A, Medela A “Dermatology Image Quality Assessment (DIQA): Artificial intelligence to ensure the clinical utility of images for remote consultations and clinical trials” Journal of the American Academy of Dermatology. 2023. doi:10.1016/j.jaad.2022.11.002
J Am Acad Dermatol. 2023;88(4):927-928
- Pearson correlation ≥0.70 with expert image quality assessment
- Real-time evaluation of focus, lighting, framing, and resolution
- Applicability to both clinical practice and clinical trial settings
DIQA is the image quality assessment algorithm that acts as a quality gate in the clinical trial workflow. It is critical for maintaining consistent image quality across investigator sites in multi-center trials.
Related validation studies
The same AI architecture and methodology used for acne scoring has been validated across multiple dermatological conditions:
Legit.Health “Automatic International Hidradenitis Suppurativa Severity Score System (AIHS4): A Novel Tool to Assess the Severity of Hidradenitis Suppurativa Using Artificial Intelligence” Published. 2025.
- Inter-observer ICC ≥ 0.727 (95% CI: 0.66–0.79) for objective severity assessment
- State-of-the-art comparison: ICC of 0.47 without the device vs. 0.727 with the device
- Same object detection + scoring methodology as acne
Validation has also been completed for APASI (psoriasis), ASCORAD (atopic dermatitis), and multiple MRMC studies (BI_2024, SAN_2024).
Regulatory-grade validation pathway
The clinical evidence follows a structured regulatory pathway:
| Standard | Scope | Application to acne scoring |
|---|---|---|
| IEC 62304 | Software lifecycle processes | The AI scoring pipeline follows a documented development lifecycle with risk-based classification |
| ISO 14971 | Risk management | Systematic risk analysis including failure modes (missed lesions, false positives, image quality) |
| IEC 62366-1 | Usability engineering | The mobile capture application has been validated for usability at investigator sites |
| MEDDEV 2.7/1 Rev 4 | Clinical evaluation | Clinical evidence compiled following the structured methodology for clinical evaluation reports |
| MDR Annex XIV | Clinical evaluation and PMCF | Post-market clinical follow-up ensures ongoing validation as the technology evolves |
Ongoing clinical validation program
| Study | Condition | Endpoints | Status |
|---|---|---|---|
| ALADIN observational study | Acne vulgaris | Lesion count, density, IGA | Published (Skin Health and Disease, 2026) |
| ALADIN post-market clinical follow-up study (DermoMedic) | Acne vulgaris | Real-world agreement with the treating dermatologist's IGA; remote monitoring | Planned; starts after MDR certification |
| DIQA validation | All conditions | Image quality assessment | Published (JAAD 2023) |
| AIHS4 validation | Hidradenitis suppurativa | IHS4 severity scoring | Published |
| MRMC study (BI_2024) | Multiple conditions | Diagnostic accuracy, severity | Completed |
| MRMC study (SAN_2024) | Multiple conditions | Diagnostic accuracy, severity | Completed |
| APASI validation | Psoriasis | PASI severity scoring | Published |
| ASCORAD validation | Atopic dermatitis | SCORAD severity scoring | Published |
For the full list of clinical evidence, see the clinical validation section.