Skip to main content

Clinical Evidence and Validation

ALADIN combines lesion count and density into an IGA-aligned severity score. Its peer-reviewed evaluation puts the results alongside expert dermatologist assessments, giving your team a transparent basis for evaluating image-based acne scoring.

How the ground truth is established

The study evaluates ALADIN and individual dermatologists against the same expert-consensus reference. This gives sponsors a clinically relevant comparison: how closely does automated scoring agree with a pooled expert assessment, and how does that compare with individual readers?

AcneSeverity-V1 supported formula optimisation and in-sample evaluation. AcneSeverity-V2 added an independent, exploratory evaluation on public-atlas images, extending the assessment to a dataset with predominantly severe acne.

Understanding the metrics

Four metrics are used to evaluate the agreement between ALADIN scores and the expert consensus. Together, they show how scores relate to the consensus and the average size of scoring differences. The displayed bands help interpret these metrics; they are descriptive guides:

Pearson correlation

The strength and direction of the linear relationship between two sets of scores. In this context: how closely ALADIN's scores track the expert consensus scores.

Perfect score: 1.0 would mean perfect linear correlation with the consensus. However, individual dermatologists themselves only achieve ~0.56–0.65 against the consensus, because acne grading is inherently subjective.

Why it matters: Shows how closely automated severity scores follow the linear pattern of expert-consensus scores across the evaluated images.

Expert comparator ▼
WeakModerateStrongVery strong
ALADIN: 0.58Expert comparator: 0.56

Spearman correlation

The strength of the monotonic (rank-order) relationship between two sets of scores. Less sensitive to outliers than Pearson. Measures whether ALADIN ranks patients in the same severity order as the consensus.

Perfect score: 1.0 would mean perfect rank agreement. Individual dermatologists achieve ~0.58 against the consensus.

Why it matters: Shows how consistently automated scoring orders the evaluated images from lower to higher severity compared with the expert consensus.

Expert comparator ▼
WeakModerateStrongVery strong
ALADIN: 0.56Expert comparator: 0.58

Cohen’s kappa (κ)

Agreement between two raters beyond what would be expected by chance. Uses quadratic weights, meaning a 2-grade disagreement is penalised more than a 1-grade disagreement. This is the primary measure of clinical agreement.

Perfect score: 1.0 would mean perfect agreement. The dermatologist results give you a clinical comparator for the automated scores.

Why it matters: Measures agreement with the expert consensus while accounting for chance and giving greater weight to larger scoring disagreements.

Expert comparator ▼
FairModerateSubstantialAlmost perfect
ALADIN: 0.53Expert comparator: 0.46

Mean absolute error

The average magnitude of scoring errors in IGA points. An MAE of 0.63 means that, on average, ALADIN’s score differs from the consensus by less than 1 IGA grade.

Perfect score: 0.0 would mean zero error. Individual dermatologists achieve MAE of 0.44–0.65 against the consensus — even experts don’t agree perfectly.

Why it matters: Puts scoring differences into familiar IGA units: the average absolute distance from expert consensus across all evaluated images. Read it alongside correlation and kappa to assess both agreement and error.

Expert comparator ▼
Lower average errorAverage error below one gradeAverage error around one gradeLarger average error
ALADIN: 0.63Expert comparator: 0.65

Severity scoring performance: ALADIN vs. dermatologists

The following results are from the peer-reviewed validation study (Sabater et al., Skin Health and Disease, 2026). They evaluate the ALADIN scoring formula, the mathematical model that converts lesion count and density into an IGA-aligned severity score. Both ALADIN and individual dermatologists are measured against the same ground truth (dermatologist consensus). Values in bold with a check mark indicate a numerically equal or more favourable point estimate: higher for correlation and kappa, lower for MAE. Check marks highlight numerical comparisons, not statistical significance.

IGA agreement

MetricMild-to-moderateSevere cases
ALADINDermatologistsALADINDermatologists
Pearson0.580.560.650.65
Spearman0.560.580.540.58
MAE0.63 (0.49)0.65 (0.63)0.64 (0.71)0.44 (0.5)
Cohen’s κ0.530.460.610.62

How to read this table

  • Agreement with expert assessment: Correlation and weighted kappa let you compare ALADIN with individual dermatologists against the same consensus. The independent dataset shows similar correlation and kappa estimates, alongside a higher device MAE.
  • Average scoring difference: MAE expresses the average absolute difference from consensus in IGA grades across all evaluated images. ALADIN has lower MAE on AcneSeverity-V1; dermatologists have lower MAE on AcneSeverity-V2.
  • Study context: V1 was used for optimisation, while V2 provides an independent, exploratory comparison. The study evaluates agreement on individual images rather than changes over a course of treatment.

Put the evidence to work for your trial

Review the scoring method, expert comparisons and image acquisition requirements together to assess how automated scoring fits your study. The published results make that assessment transparent: you can examine both agreement and error, with separate results for each dataset.

Explore the trial workflow to see how image capture, scoring and data export fit together, or review the scoring methodology for the calculation behind each endpoint.

Lesion detection performance

The following metrics are validated per the Quality Management System (IEC 62304 / ISO 14971).

  • mAP@50 (95% CI): 0.45
  • rMAE Papule (95% CI): 0.58
  • rMAE Pustule (95% CI): 0.28
  • rMAE Comedone (95% CI): 0.62
  • rMAE Nodule/Cyst (95% CI): 0.33

How to read these metrics

  • mAP@50 = 0.45 (95% CI: 0.43–0.47): The inflammatory lesion detector achieves mean average precision of 0.45 at the standard 50% IoU threshold. The acceptance criterion (≥0.21) is based on non-inferiority to published acne detection studies, where the literature range is 0.21–0.54.
  • rMAE (relative Mean Absolute Error): For per-type detection, each lesion type's counting error is compared against the inter-rater variability of expert dermatologists on the same task. The acceptance criterion is that the model's rMAE must not exceed the experts' rMAE. All lesion types pass, meaning the AI counts each lesion type at least as consistently as dermatologists disagree with each other.
Why detection precision matters less than you think

The lesion detector is not used in isolation; it feeds into the ALADIN/IGA formula, which was co-optimised with the detector. The formula constants (aa, bb) were calibrated against reference IGA scores using the detector's actual output. The publication reports end-to-end IGA agreement for the detector and formula together. This connects the component-level measurements to the severity score evaluated against expert consensus.

Additionally, the acceptance criteria are based on non-inferiority to expert inter-rater variability, a stronger validation methodology than simple accuracy benchmarks. The question is not "is the AI perfect?" but "is the AI at least as consistent as dermatologists are with each other?"

Validation datasets

DatasetImagesSourceSeverity profileAnnotatorsRole
Mild-to-moderate331Private (Clínica Dermatológica Internacional, Spain)Predominantly mild-to-moderate facial acne (IGA 0: 14, IGA 1: 65, IGA 2: 150, IGA 3: 83, IGA 4: 19)3Primary dataset for ALADIN formula design and evaluation
Severe cases27Public dermatology atlases (DermAtlas, Atlas Dermatológico)Predominantly severe acne (IGA 0: 1, IGA 1: 3, IGA 2: 10, IGA 3: 11, IGA 4: 2)2Exploratory test set for evaluating generalisation to severe cases and third-party images

The two datasets complement each other:

  • 331 images: AcneSeverity-V1, used for formula optimisation and in-sample evaluation, predominantly mild-to-moderate acne, annotated by 3 specialists
  • 27 images: AcneSeverity-V2, an independent exploratory public-atlas dataset, predominantly severe acne, annotated by 2 specialists

ALADIN peer-reviewed publication

Sabater A, et al.ALADIN: Acne Lesion And Density INdex. A Novel Tool for Automatic Acne Severity Assessment Skin Health and Disease (British Association of Dermatologists). 2026. (Provisionally accepted)

  • Development of the IGA = N^a · (D + b) formula combining lesion count and spatial density into an IGA-aligned severity score
  • Calibration of formula constants against the consensus of three board-certified dermatologists
  • Validation of the CNN-based inflammatory lesion detection model (papules, pustules, nodules)
  • Evidence that incorporating spatial density alongside lesion count improves alignment with clinical severity perception

The paper was presented as a poster at the AEDV 2025 conference in Paris prior to journal submission.

DIQA validation

Hernández Montilla I, Mac Carthy T, Aguilar A, Medela ADermatology Image Quality Assessment (DIQA): Artificial intelligence to ensure the clinical utility of images for remote consultations and clinical trials Journal of the American Academy of Dermatology. 2023. doi:10.1016/j.jaad.2022.11.002

J Am Acad Dermatol. 2023;88(4):927-928

  • Pearson correlation ≥0.70 with expert image quality assessment
  • Real-time evaluation of focus, lighting, framing, and resolution
  • Applicability to both clinical practice and clinical trial settings

DIQA is the image quality assessment algorithm that acts as a quality gate in the clinical trial workflow. It is critical for maintaining consistent image quality across investigator sites in multi-center trials.

The same AI architecture and methodology used for acne scoring has been validated across multiple dermatological conditions:

Legit.HealthAutomatic International Hidradenitis Suppurativa Severity Score System (AIHS4): A Novel Tool to Assess the Severity of Hidradenitis Suppurativa Using Artificial Intelligence Published. 2025.

  • Inter-observer ICC ≥ 0.727 (95% CI: 0.66–0.79) for objective severity assessment
  • State-of-the-art comparison: ICC of 0.47 without the device vs. 0.727 with the device
  • Same object detection + scoring methodology as acne

Validation has also been completed for APASI (psoriasis), ASCORAD (atopic dermatitis), and multiple MRMC studies (BI_2024, SAN_2024).

Regulatory-grade validation pathway

The clinical evidence follows a structured regulatory pathway:

StandardScopeApplication to acne scoring
IEC 62304Software lifecycle processesThe AI scoring pipeline follows a documented development lifecycle with risk-based classification
ISO 14971Risk managementSystematic risk analysis including failure modes (missed lesions, false positives, image quality)
IEC 62366-1Usability engineeringThe mobile capture application has been validated for usability at investigator sites
MEDDEV 2.7/1 Rev 4Clinical evaluationClinical evidence compiled following the structured methodology for clinical evaluation reports
MDR Annex XIVClinical evaluation and PMCFPost-market clinical follow-up ensures ongoing validation as the technology evolves

Ongoing clinical validation program

StudyConditionEndpointsStatus
ALADIN observational studyAcne vulgarisLesion count, density, IGACompleted; paper provisionally accepted
ALADIN clinical investigation (DermoMedic)Acne vulgarisProspective validationIn preparation (2026)
DIQA validationAll conditionsImage quality assessmentPublished (JAAD 2023)
AIHS4 validationHidradenitis suppurativaIHS4 severity scoringPublished
MRMC study (BI_2024)Multiple conditionsDiagnostic accuracy, severityCompleted
MRMC study (SAN_2024)Multiple conditionsDiagnostic accuracy, severityCompleted
APASI validationPsoriasisPASI severity scoringPublished
ASCORAD validationAtopic dermatitisSCORAD severity scoringPublished

For the full list of clinical evidence, see the clinical validation section.