Saltar al contenido principal

Known Limitations

Sponsors need to know where an AI scoring system stops, not only where it performs. This page documents the boundaries of automated CLASI component scoring in cutaneous lupus erythematosus, why each boundary exists, and how it is managed.

CLASI components not derived from an image

Three CLASI components are not observations of photographed skin, and no imaging technology produces them.

Mucous membrane lesions (CLASI-A, 0–1) require oral and nasal mucosal examination, outside the scope of a skin photography protocol.

Recent hair loss (CLASI-A, 0–1) asks whether the patient has noticed hair loss in the preceding 30 days. It is a history question with no visual correlate at a single timepoint.

Dyspigmentation duration (the CLASI-D multiplier) asks whether dyspigmentation in that patient typically persists beyond 12 months. It is a history question that doubles the dyspigmentation subtotal.

Mitigation: all three are collected through the platform as structured clinician or patient entries at the same assessment, and combined with the measured items into the final scores. They are captured, not estimated, and the assembled CLASI is complete. What the technology contributes to these three items is workflow and completeness checking rather than measurement.

Scarring, atrophy and panniculitis

The CLASI-D item combining scarring, atrophy and panniculitis is the least tractable component from surface photography. Atrophy is substantially a textural and three-dimensional finding, assessed in clinical practice partly by palpation and by how skin behaves under raking light. Panniculitis is a subcutaneous process whose surface appearance is non-specific. Mature scarring can be visually similar to surrounding skin in the absence of accompanying dyspigmentation.

Mitigation: this item is recorded by the investigator during review rather than measured, and the platform's contribution is the longitudinal image record against which the investigator judges change. Because damage accumulates over years, having a quality-controlled baseline photograph of each region to compare against is itself a material improvement over recalling the region's earlier appearance. Where the diagnosis support model is included in the configuration, it can flag panniculitis as a recognised entity, which supports the investigator's judgement without substituting for it.

Scarring compared with non-scarring alopecia

CLASI treats the two separately and for good reason: non-scarring alopecia is an activity finding that can recover, while scarring alopecia is permanent damage. They sit in different scores.

The device measures hair loss percentage, the extent of the scalp affected. It does not distinguish scarring from non-scarring loss, so the same measurement is relevant to an item in each score and resolves neither on its own.

Mitigation: the distinction is made by the investigator, who is examining the scalp in any case, and the measured percentage quantifies the extent once the type is established. The measurement removes the estimation of how much scalp is involved, which is the part of the judgement most prone to variation, and leaves the type determination with the clinician, where it belongs.

Alopecia pattern compared with alopecia extent

CLASI grades scalp alopecia by pattern: diffuse, focal in one quadrant, or focal in more than one quadrant. The device measures extent. A patient with diffuse thinning and a patient with one dense patch can present the same affected percentage and different CLASI grades.

Mitigation: hair loss percentage is reported as a continuous scalp measure alongside the clinician's pattern grade, not as a substitute for it. For studies where scalp involvement is a focus, the continuous measure is frequently the more sensitive endpoint, because it detects change within a CLASI grade that the three-point pattern item cannot resolve.

Region traceability

Every score depends on each image being attached to a patient, a visit and a named CLASI region. Without that attachment the device returns valid per-image measurements that cannot be assembled into CLASI-A, CLASI-D, or any change from baseline.

This is the most common reason a CLE image dataset cannot yield the endpoint a sponsor expected, and it is a property of the instrument rather than of the technology: a human rater given the same unlabelled images could not produce a CLASI either.

Mitigation: in a prospective study the platform supplies the metadata at capture, and the failure mode does not arise. For an existing dataset, the scope of what is achievable is established before any analysis begins, so the deliverable is agreed rather than discovered. See the trial workflow page.

Regions not photographed

CLASI is a sum over regions. A region not captured contributes nothing, and a total assembled over a reduced region set is not comparable with a whole-body total.

Mitigation: the region set is fixed at protocol design and enforced at capture, with the application tracking completion against it. Where a study deliberately restricts scope, for instance to photo-exposed sites, the reduced scope is stated in the analysis plan and the score reported against it consistently at every visit.

Extent without a physical reference is framing-dependent

An extent measured as a proportion of the photographed area carries the framing in its denominator. Two photographs of the same unchanged lesion, one taken at 30 cm and one at 20 cm, return different proportions. Nothing in the measurement itself reveals which effect is which, so a change in relative extent between two visits cannot be attributed to the patient unless the framing was held constant, and holding it constant by eye across years, sites and staff is not a realistic assumption.

This is a property of the quantity, not a defect in the model, and it applies equally to a proportion estimated by a human observer.

Worse, the denominator is often not the photograph either. Several measurements are expressed as a proportion of the skin the model could see, which is itself a segmentation: hair, clothing, dressings and camera angle all change it. A scalp photograph in which hair covers most of the head yields a small visible-skin area, so a lesion occupying much of that exposed skin returns a high proportion even though it covers a modest part of the scalp. The figure is arithmetically correct and means something narrower than it appears, and two visits with different hair positioning are not comparable on it.

Mitigation: capture with a calibration marker in frame and report absolute area in square millimetres, which is independent of both working distance and how much skin happened to be visible. Where a study cannot use markers, for instance in a retrospective analysis of images already captured, relative extent should be reported as a descriptive measure of each image, with its denominator stated, and not used as a change-from-baseline endpoint. Intensity measurements such as erythema grade are unaffected by framing and remain usable in both cases.

Hyperpigmentation is not separately measured

CLASI-D scores dyspigmentation as one item covering both loss and excess of pigment. The device measures the depigmented component as a continuous extent. It does not produce an equivalent measurement of post-inflammatory hyperpigmentation.

The consequence is that CLASI-D dyspigmentation is partly measured and partly judged. Where a region shows depigmentation, the measurement satisfies the item and supports the continuous analysis described above. Where the pigmentary change is hyperpigmentation alone, which is common in darker phototypes and in resolving lesions, the item remains a clinician observation and no continuous measure is available for it.

Mitigation: the item is completed by the investigator at review, as the other non-measured components are, so the assembled CLASI-D is complete. What is unavailable is the sub-CLASI resolution on the hyperpigmented component specifically. Studies whose endpoint depends on quantifying post-inflammatory hyperpigmentation should treat that as out of scope for automated measurement today.

Skin phototype, and why it runs in both directions

Phototype is usually discussed as though it degrades every measurement in the same direction. For CLE it does not, and the difference is worth stating precisely because it changes which measurements a study can rely on in which population.

Erythema is harder in darker skin. Redness is judged against surrounding skin, and in deeply pigmented skin it shows lower contrast and shifts toward violaceous rather than red. This is a recognised weakness of visual severity instruments generally, and it affects human raters and automated measurement alike.

Depigmentation is harder in lighter skin, for exactly the same reason inverted. A depigmented plaque on deeply pigmented skin is a high-contrast target and is measured well. On a light phototype there is little contrast between a depigmented lesion and the surrounding skin, and a segmentation model has correspondingly little to work with.

The practical rule is that contrast, not phototype, is what governs reliability, and the two signs put their difficult cases at opposite ends of the phototype range. A study in a phototype-diverse population should expect its erythema and its pigmentation measurements to be reliable in different subgroups, and should power and interpret them accordingly rather than assuming a single direction of bias.

Mitigation: colour calibration markers materially reduce the illumination component of this variability and are recommended for any CLE study in a phototype-diverse population. Where a pigmentation endpoint is central and the population is predominantly light-phototype, the measurement should be treated as descriptive and confirmed against clinician assessment. Performance across skin types is tracked as a standing question in post-market surveillance; see skin type performance.

Low-contrast lesions

Following from the above, the general form of the constraint is that every extent measurement depends on the lesion being distinguishable from the skin around it. A lesion whose colour closely matches surrounding skin, whether because of phototype, because it is resolving, or because the illumination flattened the difference, is measured less reliably than a high-contrast one, and the failure is silent: a mask is still returned.

Mitigation: masks are returned with every assessment and are reviewable, so a measurement can be checked against what was actually segmented rather than taken on trust. Where extent is an endpoint, the protocol should include review of the segmentation overlays, and DIQA plus calibration markers should be used to hold illumination stable so that contrast reflects the lesion rather than the lighting.

CLASI assembly is study configuration

The device measures clinical signs and exposes a set of established scoring systems. CLASI is not one of its built-in presets, and it is not a score that appears by pointing the technology at a CLE patient.

Assembling CLASI-A and CLASI-D from the measured components is configuration work carried out during study setup: fixing the region set, mapping each sign measurement onto its CLASI item scale, defining the thresholds at which an extent measurement satisfies a present-or-absent item, and wiring in the three clinician-entered components. None of that is research, and all of it has to be specified, agreed and frozen before the first patient is imaged.

Mitigation: the configuration is a defined deliverable of study setup rather than an open question, and it is recorded in the study configuration so that a value means the same thing at the final visit as at baseline. The practical consequence is on timelines rather than feasibility: a CLE study needs this specification agreed during protocol design, alongside the imaging protocol, not after sites are live.

Validation status in cutaneous lupus erythematosus

This is the limitation that matters most for endpoint planning, and it should be stated plainly.

The sign measurements described in this section are components of a CE-marked medical device, developed and validated within the indications they were built for. Erythema, desquamation and induration were developed and validated in inflammatory dermatoses; hair loss measurement in alopecia; diagnosis support across a broad set of conditions that includes cutaneous lupus erythematosus.

No CLASI-specific validation study in a CLE population has been performed. The signs CLASI is built from are measured by the device, and the agreement between assembled CLASI-A and CLASI-D and expert clinician CLASI scoring in CLE patients is not yet characterised. Neither is the performance of pigmentation measurement in this specific population.

Mitigation: this is a study to run, not a defect to work around, and it is the natural first step in a CLE programme. A validation study comparing automated component measurement against expert CLASI scoring in a CLE cohort establishes agreement, quantifies reliability against the manual inter-rater benchmark, and produces the evidence a regulator would expect before an AI-derived CLASI is proposed as anything beyond an exploratory endpoint.

The substrate for that study frequently already exists, and it is worth naming precisely. Any sponsor that has run a CLE trial holds a longitudinal image set whose visits already carry clinician-scored CLASI-A and CLASI-D. That pairing, images on one side and expert scores on the other, is exactly what a validation needs, and it means the first study can be retrospective: no new patients, no new sites, no new capture. Its yield depends on how well the images are labelled by patient, visit and region, which is the constraint described above, and establishing that is a scoping exercise measured in days rather than a study.

A retrospective validation on an existing scored dataset is therefore the cheapest available step, and it is what de-risks the design of any prospective work that follows.

Until that evidence exists, an AI-derived CLASI belongs in a protocol as an exploratory or supportive measure alongside clinician CLASI, which is also the configuration that generates the paired data the validation needs.

Cross-cutting limitations

The following limitations apply to all indications scored by the platform, not only cutaneous lupus erythematosus.

Photograph-based assessment

The AI analyses clinical photographs, not live patients. Certain clinical features that require palpation (e.g., induration, or plaque thickness) or observation under specific conditions are estimated from visual cues only. This is an inherent limitation of any remote or image-based assessment method.

Mitigation: The imaging protocol standardises capture conditions (lighting, distance, angle), and the DIQA quality gate rejects images that do not meet minimum quality standards for focus, lighting, framing, and resolution. The acceptance criterion for each AI model is non-inferiority to expert inter-rater variability on the same photographs, ensuring the AI is at least as consistent as dermatologists working from the same modality.

Fitzpatrick skin type V–VI performance

Performance is lower for darker skin types due to the global underrepresentation of Fitzpatrick V–VI skin in dermatology image datasets. This is an industry-wide challenge that affects both AI systems and human assessors.

Mitigation: Stratified performance metrics are published transparently (see Performance Across Skin Types). Active dataset diversification is ongoing through targeted data sourcing (DDI, SkinDeep, Full Spectrum Dermatology) and post-market clinical follow-up (PMCF) monitoring. All Fitzpatrick groups currently exceed minimum acceptance thresholds.

Subjective ground truth

The reference standard for severity scoring is the mathematical consensus of multiple expert dermatologists — not an objective measurement. The AI cannot be “more correct” than the experts it was trained against.

Mitigation: This is not a limitation of the AI specifically, but of the clinical assessment itself. The same constraint applies to any human rater. The consensus of 2–3 independent experts is the best available approximation of truth and the same standard used by the FDA and EMA for reference standards in dermatology clinical trials. The AI matching this consensus represents the realistic ceiling of performance.

Model version specificity

All performance metrics reported in this documentation apply to a specific validated model version. Model updates — including retraining, architecture changes, or threshold adjustments — require full re-validation per IEC 62304 before deployment.

Mitigation: The model version is locked at study initiation. No mid-study model updates occur. This ensures that every patient in a trial is scored by the same model, preserving endpoint integrity throughout the study.

Decision support, not autonomous diagnosis

The system provides severity scoring to support clinical decisions. It does not replace clinical judgement, and it does not make autonomous diagnostic or treatment decisions. All AI-generated scores should be interpreted by qualified healthcare professionals within the context of the patient’s overall clinical presentation.

How limitations are managed

All limitations documented on this page are tracked within the formal risk management process (ISO 14971) and the software development lifecycle (IEC 62304). Each has been assessed for clinical risk, and mitigations implemented where the residual risk is not already acceptable.

The post-market clinical follow-up programme under MDR Annex XIV continuously monitors real-world performance. A new limitation identified through post-market surveillance triggers a formal risk assessment and, where necessary, corrective action.