Aller au contenu principal

Known Limitations

Clinical trial sponsors need to understand not just what an AI scoring system can do, but where its boundaries are. This page documents the known limitations of the automated IHS4 scoring provided by Legit.Health, explains why each limitation exists, and describes how it is managed. Transparency about limitations is a regulatory expectation (ISO 14971, EU AI Act Article 13) and a prerequisite for informed protocol design.

A physical examination is always required​

Hidradenitis suppurativa cannot be assessed from photographs alone: much of the disease develops beneath the skin, and properties such as the fluctuance of an abscess are only confirmed by touch. Deep nodules, subcutaneous tracts and the extent of tunnels under the skin cannot be assessed from surface photography, so the AI scores only the lesions visible at the surface. Its detections are therefore a first screening, never the final assessment.

Mitigation: Every assessment includes the investigator review, after which the counts and scores are recalculated. For studies that require a formal assessment of deep involvement, the protocol specifies that palpation or ultrasound findings are recorded separately.

Each photograph is analysed on its own​

The AI analyses every photograph independently and does not merge the detections of several photographs into one. A lesion too large to fit in one photograph, or one that appears in two overlapping photographs, is detected separately in each.

Mitigation: The imaging protocol asks for each lesion to be captured whole in a single photograph where possible, and the investigator review is where the detections of all the photographs of a visit are checked before the counts are confirmed.

Draining and non-draining tunnels are counted together​

Hidradenitis suppurativa develops mostly beneath the skin, as interconnected tunnels and cavities in the subcutaneous tissue. A photograph shows the surface: a tunnel opening, the tract it outlines, and any discharge visible at the moment of capture. Whether a tunnel is actively draining depends on information the image does not carry, such as the patient's recent history, drainage between visits and what the clinician finds on examination. The AI therefore detects tunnels without separating draining from non-draining ones, and counts every visible tunnel in the IHS4 tunnel term (×4), the highest-weighted lesion in the score.

Mitigation: The tunnel count is reported as its own field and is confirmed by the investigator during review, so a protocol that needs the draining status can record it at the visit alongside the count. Because the count is made the same way at every visit, change from baseline stays consistent even where the absolute tunnel count differs from a clinician's draining-only count.

Abscess vs. nodule differentiation​

Distinguishing abscesses from inflammatory nodules is difficult even for experienced clinicians. Abscesses are characterised by fluctuance, which is a tactile property, and by visual features such as surrounding erythema and an irregular surface, but the boundary between a large nodule and a small abscess is often ambiguous.

Mitigation: This ambiguity is a characteristic of HS lesion assessment itself, not specific to AI, and it is the main reason manual IHS4 has low inter-rater agreement. The model is trained on images annotated by a board of specialists and unified into a consensus reference, so it applies one consistent boundary at every site and every visit. Where touch decides, the investigator corrects the type during review. See Clinical evidence.

Only photographed regions contribute to the score​

Hidradenitis suppurativa is photographed region by region, and only where there are lesions: there is no full-body capture set. A region that is not photographed cannot contribute to the counts, so an affected region missed at capture lowers the IHS4 for that visit.

Mitigation: The imaging protocol lists the regions to examine at every visit, and the capture is performed at the site by the study staff as part of the clinical examination, so every affected region is identified before it is photographed. The DIQA quality gate checks each image before submission.

Scar tissue and post-inflammatory changes​

Chronic HS can leave extensive scarring and post-inflammatory changes. The AI must distinguish active lesions, which contribute to IHS4, from inactive scars, which do not. Dense scar tissue can also obscure new lesions, so a region with heavy scarring is harder for the algorithm to read than one without.

Mitigation: The model is trained to differentiate active inflammatory lesions from post-inflammatory scars. In areas of very dense scarring with superimposed active disease, some detection difficulty may persist, and patients with extensive scarring should be discussed during protocol design. Longitudinal tracking helps: baseline scarring is documented and changes from baseline reflect active disease.

Cross-cutting limitations​

The following limitations apply to all indications scored by the platform, not just hidradenitis suppurativa.

Photograph-based assessment

The AI analyses clinical photographs, not live patients. Certain clinical features that require palpation (e.g., induration, or plaque thickness) or observation under specific conditions are estimated from visual cues only. This is an inherent limitation of any remote or image-based assessment method.

Mitigation: The imaging protocol standardises capture conditions (lighting, distance, angle), and the DIQA quality gate rejects images that do not meet minimum quality standards for focus, lighting, framing, and resolution. The acceptance criterion for each AI model is non-inferiority to expert inter-rater variability on the same photographs, ensuring the AI is at least as consistent as dermatologists working from the same modality.

Fitzpatrick skin type V–VI performance

Performance is lower for darker skin types due to the global underrepresentation of Fitzpatrick V–VI skin in dermatology image datasets. This is an industry-wide challenge that affects both AI systems and human assessors.

Mitigation: Stratified performance metrics are published transparently (see Performance Across Skin Types). Active dataset diversification is ongoing through targeted data sourcing (DDI, SkinDeep, Full Spectrum Dermatology) and post-market clinical follow-up (PMCF) monitoring. All Fitzpatrick groups currently exceed minimum acceptance thresholds.

Subjective ground truth

The reference standard for severity scoring is the mathematical consensus of multiple expert dermatologists — not an objective measurement. The AI cannot be “more correct” than the experts it was trained against.

Mitigation: This is not a limitation of the AI specifically, but of the clinical assessment itself. The same constraint applies to any human rater. The consensus of 2–3 independent experts is the best available approximation of truth and the same standard used by the FDA and EMA for reference standards in dermatology clinical trials. The AI matching this consensus represents the realistic ceiling of performance.

Model version specificity

All performance metrics reported in this documentation apply to a specific validated model version. Model updates — including retraining, architecture changes, or threshold adjustments — require full re-validation per IEC 62304 before deployment.

Mitigation: The model version is locked at study initiation. No mid-study model updates occur. This ensures that every patient in a trial is scored by the same model, preserving endpoint integrity throughout the study.

Decision support, not autonomous diagnosis

The system provides severity scoring to support clinical decisions. It does not replace clinical judgement, and it does not make autonomous diagnostic or treatment decisions. All AI-generated scores should be interpreted by qualified healthcare professionals within the context of the patient’s overall clinical presentation.

How limitations are managed​

All limitations documented on this page are tracked within the formal risk management process (ISO 14971) and the software development lifecycle (IEC 62304). Each limitation has been assessed for clinical risk, and mitigations have been implemented where the residual risk is not already acceptable.

The post-market clinical follow-up (PMCF) programme under MDR Annex XIV continuously monitors real-world performance. Any new limitation identified through post-market surveillance triggers a formal risk assessment and, if necessary, corrective action.