Cutaneous Lupus Erythematosus Severity Endpoints for Clinical Trials
The AI scoring provided by Legit.Health automates the visual components of the CLASI (Cutaneous Lupus Erythematosus Disease Area and Severity Index), the reference outcome measure in cutaneous lupus erythematosus (CLE) drug development.
CLASI is unusual among dermatological indices because it produces two independent scores from the same examination: an activity score that responds to treatment within weeks, and a damage score that accumulates irreversibly. A trial that reports only one of them is reporting half the disease.
Why automated CLASI for clinical trials?
Manual scoring
Same patient, same image — three dermatologists, three different scores:
Typical inter-rater agreement: Cohen's κ = 0.41–0.60
AI scoring
Same patient, same image — always the same result:
Intra-rater variability: 0.00 — perfectly reproducible
Inter-rater variability
Different investigators assign different scores to the same patient
Slow and costly
Lesion counting takes 5\u201315 minutes per patient and is impractical at scale
Inflated sample sizes
Scoring noise masks treatment effects, requiring larger enrolments
Site training burden
Calibration exercises are costly, time-consuming, and imperfect
CLASI compounds the general problem in three specific ways.
The instrument is long. Two activity items and two damage items are scored in each of 13 anatomical regions, plus scalp alopecia and scalp scarring by quadrant and two patient-level questions. A single complete assessment is over fifty judgements, which is why CLASI is well established in trials and thinly used in routine care.
The hardest item is the primary one. CLASI activity is dominated by erythema, graded on a four-point scale that runs from faint pink through red to dark red or violaceous. Grading colour by eye is exactly the task in which observers differ most, and it differs again under a different examination-room light.
Activity and damage are easy to confuse. Residual erythema in a healing lesion and early dyspigmentation in a damaged one can look similar at a single visit. Separating them reliably is what makes the two scores independent, and it is the separation most vulnerable to rater drift over a long study.
The two scores
Erythema and scale/hypertrophy per region, plus alopecia and mucosal involvement
Responds to treatment; the usual efficacy endpoint
Dyspigmentation and scarring/atrophy per region, plus scalp scarring
Accumulates irreversibly; the long-term burden measure
What the AI measures
The device quantifies the visual signs that CLASI is built from, each from a standard smartphone photograph of a single anatomical region.
| Measurement | Output | CLASI component it serves |
|---|---|---|
| Erythema intensity | Graded severity of redness | CLASI-A erythema |
| Erythema extent | Affected area, relative or absolute | Lesion burden within the region |
| Desquamation intensity | Graded severity of scaling | CLASI-A scale |
| Induration intensity | Graded plaque thickening | CLASI-A hypertrophy |
| Depigmentation extent | Affected area, relative or absolute | CLASI-D dyspigmentation |
| Lesion area | Physical area in mm², with marker capture | Longitudinal lesion measurement |
| Hair loss percentage | Affected proportion of the scalp | CLASI-A and CLASI-D scalp items |
| Diagnosis support | Ranked differential including CLE | Screening and eligibility review |
| DIQA image quality | Quality score with accept or recapture | Applies to every capture |
| Skin and body segmentation | Pixel-level masks | Underlies every extent figure |
CLASI-D collapses every pigmentary change into one present-or-absent item. The device measures the depigmented component as a continuous extent instead, which carries more information than the instrument requires and is what makes the repigmentation measure below possible. Post-inflammatory hyperpigmentation is not separately quantified, so CLASI-D dyspigmentation is partly measured and partly a clinician judgement; the limitations page sets out the boundary.
Extent needs a physical reference to be comparable over time
An extent figure can be expressed two ways, and the difference decides whether it is usable as a longitudinal endpoint.
Relative extent is the affected proportion of the photographed area. It needs nothing beyond the photograph, and it is only comparable between two images framed the same way. A photographer standing closer at week 12 than at baseline reduces the denominator, and the proportion moves without anything having changed in the patient.
Absolute area is the lesion measured in square millimetres. It requires a physical size reference in the frame, supplied by the calibration marker, and it is independent of how far away the camera was. It is therefore the form that supports change from baseline.
For a study where lesion extent or pigmentation is an endpoint rather than a descriptive measure, capture with markers and report absolute area. See calibration markers.
Full item-by-item mapping, including the CLASI components that are not derived from an image, is on the scoring methodology page. Boundaries of the technology are documented on the limitations page.
Repigmentation: an outcome CLASI cannot express
CLASI-D records dyspigmentation as present or absent in each region, and damage is designed to accumulate rather than resolve. Between those two properties, the instrument has no way to represent a patient whose depigmented plaques are partially repigmenting, which is a change patients notice and value, and one that a therapy may produce well before anything else moves.
Because the device measures hypopigmentation as a continuous extent rather than a flag, the change in that extent across visits is directly reportable:
- Repigmentation trajectory: reduction in hypopigmented area from baseline, per lesion and per region
- Partial response resolution: a lesion repigmenting from its periphery inward registers as a continuous change, where CLASI-D holds at the same value until the item flips
- Separated from activity: because erythema is measured independently, repigmentation is not confounded by residual inflammation in the same lesion
This is an exploratory endpoint rather than a CLASI item, and it is the clearest example of the general point: the measured components carry more resolution than the instrument they feed, and a protocol can use both.
Region-level scoring is a data requirement, not a detail
CLASI is not a whole-body score that happens to be broken into parts. It is a sum of independently scored anatomical regions, and the arithmetic only exists once each image is attached to a patient, a visit and a named region.
A set of photographs without that attachment yields per-image sign values and nothing else: no CLASI-A, no CLASI-D, no change from baseline. This is the single most consequential design decision in a CLE imaging study, and it is settled at protocol design rather than at analysis. The imaging protocol page specifies the capture and metadata requirements that make region-level scoring possible.
Longitudinal severity tracking
Because activity and damage move on different timescales, CLE studies benefit particularly from automated longitudinal tracking. From the first follow-up visit onward the platform reports:
- CLASI-A trajectory: absolute and percentage change from baseline at each visit, the usual basis for a 50% activity response endpoint
- CLASI-D trajectory: accumulation over the study, which should be flat under an effective therapy
- Target lesion trajectory: per-lesion area, erythema and pigmentation for the designated target lesion, followed at every visit
- Repigmentation: reduction in hypopigmented area from baseline, per lesion and per region
- Per-sign trends: whether an improving activity score is driven by erythema, scale or both
- Per-region trends: which anatomical regions are responding and which are not
- Protocol adherence: whether assessments are captured at the scheduled intervals

Because the same image scored twice returns the same values, a change between two visits is a change in the patient rather than a change in the observer. Over a multi-year study with site turnover, this is the property that keeps a damage score interpretable.
The platform provides automated severity scoring as decision support. Scores are interpreted by qualified healthcare professionals within the patient's overall clinical context; the platform does not replace clinical judgement or make autonomous diagnostic decisions.
Regulatory status
Software lifecycle per IEC 62304, usability per IEC 62366-1, risk management per ISO 14971, clinical evaluation per MEDDEV 2.7/1 Rev 4 and MDR Annex XIV.