Considerations for image-based assessment
Every study that assesses skin from photographs meets the same set of questions, whatever the indication and whatever the camera. What can be derived from one photograph, and what needs several read together. How measurements from several views combine into one number for a lesion, a region or a subject. What has to be decided before the first image is captured, because it cannot be recovered afterwards.
This page states those considerations once, in general terms. It deliberately carries no dimensions, distances or protocol counts: each of those belongs to the page that owns it, which is the indication's imaging protocol for the view set, and the "Calibration Markers" section for anything to do with the physical reference and the camera.
The content of the image is the boundary of the assessment. Whatever is not captured within the frame cannot be evaluated, and this limitation is a property of the input rather than of the reader: it constrains the clinician and the Legit.Health platform identically. It follows that the questions addressed below are matters of protocol design, not of reader selection.
What the photograph contains is what can be assessed
Ask a dermatologist to measure an ulcer that has been captured as several overlapping photographs, and they cannot produce a single figure by looking harder. What they do is read each photograph, obtaining several readings, and then apply a rule: add them where the photographs show different lesions, or take the one photograph in which the lesion is complete where they show the same lesion from different angles.
The platform behaves the same way, and for the same reason. It reads each image and returns a result for it. Neither reader can assess what the frame does not show, and neither invents information the photographs do not carry.
There is one genuine asymmetry, and it is worth naming plainly rather than leaving it implicit. A clinician standing in front of the subject can reposition, palpate, and lay a tape across a lesion that no single photograph frames. That is a capability of being in the room, not sharper perception, and it is absent the moment the assessment is made from an image, which is the situation in any photographic or decentralised trial. What runs the other way is repeatability: the same image yields the same number every time, across visits, sites and readers, whereas human readings of the same photograph vary between readers and within the same reader over time.
The unit of analysis: one image, or a set read together
Every output has a unit of analysis, and stating it per output is the first thing a study fixes.
Image level: the output is derived from a single photograph, and one result is returned for each image, keyed by the identifier that image arrives with. Most visible-sign assessments, single-lesion measurements and per-photograph lesion counts sit here.
Image-set level: the output cannot be derived from one photograph by definition, so a defined set of images is the unit and one result is returned per set. Any score depending on an automatic calculation of body surface area is the clearest case, because the denominator is the body rather than the frame. Where a set is the unit, an incomplete set is an incomplete input, and assembling the set correctly matters as much as capturing each image in it.
This holds identically for counts and for areas. A lesion count over an anatomical region is no more contained in one photograph than an area is, so a count that has to cover a region needs the region's view set defined just as precisely.
Adding readings across views, and when that is wrong
Once each photograph has been read, combining the readings is arithmetic, and the arithmetic is only valid in one of the two cases below. The failure mode in the other case catches a human reader exactly as it catches an automated one.
Double counting is not a subtle risk. Two overlapping views of one inflamed nodule are two photographs each containing that nodule, so a naive total reports it twice, and the score built on that total moves for a reason that has nothing to do with the subject.
A defined protocol is what makes composition possible
The distinction that matters is not between a human reader and an automated one. It is between a defined view set and an ad-hoc one.
Where the protocol fixes the perspectives in advance, the relationship between them is known, and the platform composes across them and resolves the overlap as part of the score:
- In acne, the perspectives are fixed by the protocol, and lesions appearing in more than one of them are reconciled against detected facial landmarks so that each lesion is counted once. See the acne "Scoring Methodology" page, under per-perspective scoring and overlap handling.
- In psoriasis, each image is segmented into the body regions the score is defined over, so pixels are attributed to a region rather than to a photograph. See the psoriasis "Scoring Methodology" page, under body region segmentation.
- In hidradenitis suppurativa, counts are aggregated per anatomical region and then globally into the weighted score. See the hidradenitis suppurativa "Scoring Methodology" page, under count aggregation.
- In alopecia, the scalp is captured as distinct regions with no meaningful overlap, and the total is a weighted sum of the regional results. See the alopecia "Scoring Methodology" page, under counting methodology.
Where the view set is not defined in advance, and photographs simply arrive as an undeclared collection of angles on the same lesion, no reader can compose them into one measurement, because nothing in the images states how they relate to each other. That is the honest limit, and it is a property of the input rather than of the reader. It is also the reason the answer to "can you measure this complex case?" is usually a question back: define the views, and the outputs that compose across them follow.
Three cases that cover most of what arrives
Nearly every capture question a study brings us is one of three cases, and they are worth setting out side by side, because what changes between them is the unit of analysis and therefore what has to be fixed in advance.
One lesion in one photograph
Several lesions in one photograph
An affected area larger than one frame
| The case | What the frame holds | Unit of analysis | What composes across images | Fix in advance |
|---|---|---|---|---|
| One lesion | The lesion and the skin around it | The image | Nothing has to | Which lesion, and reproducing the view at later visits |
| Several lesions | More than one distinct lesion | The image, with a result per lesion in it | Counts and the areas of distinct lesions, because they do not overlap | Which lesion is the target, if an endpoint follows one |
| A large affected area | Part of a region that continues past the frame | The defined set of images covering the region | Regional totals, where the views tile the region without overlapping | The view set for the region, and whether the output needs the whole region |
One lesion in one photograph
The straightforward case, and the one the other two are measured against. The photograph holds the lesion and the skin around it, one result is returned for that image, and where the output is a real-world measurement the frame also carries a physical reference. Nothing has to compose, so nothing has to be declared beyond which lesion is being followed.
What still has to be got right is repetition. A lesion photographed at a different angle or a different distance at a later visit is comparable to itself only by chance, and change over time is usually the endpoint. That is a capture discipline rather than an analysis question, and it is the same discipline a photographic archive needs in order to be read consistently by a person.
Several lesions in one photograph
A frame holding several lesions is not a complication. For a count it is the intended input, and counts within one frame are additive without qualification, because distinct lesions do not overlap each other. The areas of distinct lesions in the same frame add for the same reason. This is the sum a clinician performs, and it is arithmetically sound for either reader.
Three things do have to be decided, and all of them are about the protocol rather than the reader:
- Which lesion is the target, where an endpoint follows one. A photograph of several lesions does not state which of them the study is tracking. The protocol states it, typically by identifying the target on a body diagram when it is selected and then capturing that same lesion at every subsequent visit. A clinician handed the same photograph without that diagram has exactly the same difficulty, and guesses no better.
- One photograph per lesion, where a per-lesion result is wanted. Where the study wants a result for a representative lesion, and separately for others, the input that gives it is one photograph per lesion, each framed with that lesion as the principal element and with as few other lesions in it as the anatomy allows. Each photograph is then read on its own and returns its own result, so separate lesions produce separate analyses without anything having to be disentangled afterwards. A single wide frame containing all of them returns one result for the frame, which is the right input for a count and the wrong one for a per-lesion endpoint.
- The framing trade-off. A frame wide enough to hold several lesions puts less detail on each of them, and where the output is a real-world measurement the physical reference in that frame has to serve all of them. So where the endpoint is a per-lesion measurement, a frame per lesion is the better input, and where the endpoint is a count over a region, the wider frame is the right one. The practical placement guidance is in the "Image Capture" section.
Nothing identifies the target on the photograph itself. An arrow, a circle, a highlight or a written label added to an image is read as part of that image rather than as an instruction about it, and it can change the result of any output that reads the area it covers, so the way to point at one lesion among several is a separate photograph of it, not a mark on a shared one. Identification belongs on the body diagram and in the labels the image arrives with, and this is the same reason set out in the "What Arrives with the Image" section below.
The same photograph can therefore be an excellent input for one output and a poor input for another, which is why naming the output comes before discussing the images.
An affected area larger than one frame
Here the region rather than the lesion becomes the subject, and the answer follows the distinction already made above rather than adding a new one:
- Where the protocol tiles the region into defined views that do not overlap, each view is read and the results compose into a regional total. This is how region-based scores are built, and it is the case the indication protocols are designed around.
- Where the views overlap, or arrive as an undeclared set of angles, they cannot be summed, because the overlap would be counted twice. Either the protocol defines the views so the overlap is known, or a single complete view is used.
- Where the output is a proportion of the region affected, a photograph of part of the region cannot carry it at all. The denominator is the region, and a frame that holds only part of it does not contain the denominator. This is an output that is image-set level by definition, so an incomplete set leaves the affected-surface part of a score empty while the visible-sign part is still returned from the close-ups that are present.
Underneath all three, the rule the platform follows is the rule a professional reviewing the images follows, and it is worth stating arithmetically because it is what determines who does the remaining work. One image is one input and yields one output. Two images are two inputs and yield two outputs, and those two outputs inherit whatever relationship the two inputs had: if the frames overlap, the outputs overlap, and if the frames do not overlap, neither do the outputs. Nothing in the second case has been added or removed by the analysis; it simply reports what each frame contains.
Reconciling those outputs into one figure for the region is therefore a post-processing step rather than part of reading the images, and unless it has been agreed in advance for a given output it is outside the scope of the analysis we return. Where it is agreed, it is agreed because the protocol fixed the perspectives, which is exactly what the indication protocols do in the examples above.
That is the reason imaging protocols are written around standardised perspectives rather than around a number of photographs. A standardised perspective set makes the relationship between the frames known before any of them is captured, so the reconciliation is decided once in the protocol instead of being attempted afterwards on images that no longer state how they relate to each other.
The human parallel is at its clearest here. A clinician judging how much of a region is involved works against the whole region, using a rule of thumb such as the patient's palm as a proportion of their total body surface. From a close-up of part of the area, neither reader can recover the proportion, because the information is not in the photograph.
One practical consequence, and it belongs to the camera rather than to the analysis: with a fixed working distance the operator cannot simply step back to fit more in, and the physical reference still has to be fully inside each frame. Both are covered in the "Calibration Markers" section, and both are reasons to raise a large affected area before the first visit rather than at the first data transfer.
Standardisation and view labelling
The labels that arrive with the images are what carry the meaning that makes any of the above possible. A filename and an anatomical view label are what distinguish several lesions from several views of one lesion. Without them, a total cannot be assembled by a clinician or by the platform, because both would be combining numbers whose relationship is unknown.
Two things follow for the imaging procedure a study writes:
- Name the view, not just the image. Each photograph should carry the anatomical view it belongs to, and where a region is covered by several photographs, the fact that they belong to one region should be visible in the labelling rather than inferred later.
- Keep the protocol identical across visits. A view captured differently at a later visit is comparable to itself only by chance, and change over time is what most trials are measuring. The per-indication imaging protocols set out the view sets, for example the standard anatomical region protocol in the hidradenitis suppurativa "Imaging Protocol" page, and the practical capture guidance sits in the "Image Capture" section.
Where an output is a real-world measurement rather than a graded appearance, that image also needs a physical reference inside the frame, which is what the "Calibration Markers" section covers. Counts and graded visible signs do not need one. This is the distinction to draw early when the requirement is stated as measuring everything on every photograph.
What arrives with the image
The analysis reads the image as it is received, so anything added to a photograph after capture is part of what is read. A de-identification box, a mask or a frame, a mark, drawing or annotation, an overlay, or a crop is, to the models, neither skin nor lesion: it is part of the image. It can change the result of any output that reads the area it covers, and the effect is not confined to that area, because several outputs are derived from the image as a whole.
The optimal input is the original image, as it came off the camera, at its native resolution and with nothing added, removed or covered. Every processing step applied between capture and transfer can only subtract information: resizing and recompression remove detail the models use, a rotation or a crop changes what the frame contains, and an overlay replaces skin with something that is not skin. None of it can be recovered later, so where a study can transfer the original alongside any processed copy, the original is the version to analyse.
Anonymisation and its effect on the analysis
Anonymisation is the most common reason an image is altered before it is transferred, and it deserves stating plainly rather than being left as a technicality. Any part of the image that is blurred, cropped or covered, a black box most often, reduces the usable information in the frame. The more of the lesion or of the surrounding skin that is obscured, the larger the potential effect on the result. This may read as obvious, and it is exactly the point: the model can only assess what it sees, so where a large part of the image is hidden or missing, the quality of the assessment falls with it.
The parallel with a clinician holds here as everywhere else on this page. A doctor asked to evaluate a lesion with part of it hidden cannot judge its size, shape or characteristics accurately either, and would say so rather than estimate. The limitation belongs to the input, not to the reader.
What follows is within a sponsor's control, because it concerns what happens to an image between capture and transfer:
- Keep any covered or blurred area as small as possible, and above all keep it away from the lesion and the skin immediately around it.
- Confirm with the clinical team whether a less obtrusive method is acceptable, for example cropping outside the lesion area rather than covering anything inside it.
- Capture well in the first place, since a clear, well-lit and focused photograph is required whatever the anonymisation procedure. The practical guidance for this is in the "Image Capture" section.
Preserving as much of the visual information as possible is the single most useful thing a site can do for the accuracy of the results it gets back.
Missing and non-conforming inputs
Images that do not satisfy the required inputs fall outside the analysis for the affected subject, timepoint or set. Nothing is substituted, interpolated or inferred in their place, and a result row is returned for every photograph transferred, so what fell outside the analysis is visible in the delivery rather than silently absent. A body region for which no conforming photograph exists is reported as absent rather than scored, and that does not invalidate the regions that are present.
Where the protocol defines it, an incomplete set degrades gracefully instead of failing. In psoriasis, when part of the standardised set is missing, visual signs are still scored from the close-ups that are available even though the surface-area-dependent inputs are not, which is described under local PASI from visual signs in the psoriasis "Imaging Protocol" page. Graceful degradation is a designed behaviour per output, not a general guarantee, so it is confirmed for the outputs a given study depends on.
Decisions to fix before first capture
Each of these is cheap to decide in advance and expensive to decide afterwards, because an archive cannot be re-captured once a subject has left the clinic. Together they make a good agenda for an output alignment meeting.
| Decision | Why it cannot wait |
|---|---|
| The output measures wanted per cohort | Everything else depends on it. The imaging protocol, the view set and the physical reference are all consequences of which measures are wanted, so nothing downstream can be settled while this is open. |
| The unit of analysis per output | Determines whether results arrive per image or per set, and therefore how the analysis dataset is shaped and how endpoints are computed. |
| The view set, and the labels that identify it | The labels are what make aggregation possible at all. They have to be decided before capture, because a photograph cannot be relabelled with information nobody recorded. |
| Which lesion each endpoint follows | Where an endpoint tracks one lesion among several, the target is identified when it is selected and captured again at every visit. A photograph cannot be asked afterwards which lesion mattered. |
| The rule for combining readings across views | Whether readings are added, or one view is selected, or the score composes across a defined perspective set. Deciding it in advance is what stops a plausible but wrong total from being computed later. |
| A physical reference where a measure is metric | Applies only to real-world measurements. It affects the images themselves, so it cannot be added retrospectively to an archive captured without it. |
| The capture system and its confirmation | Both the reference size and the framing follow from the camera and how it is used, and confirming them takes lead time before the first visit. |
Sponsors and imaging partners working through these with us usually find the conversation short, because most of it is deciding what to write down rather than what to build. Where you are planning a study and want these settled early, the "Clinical Trial Experience" page describes how a deployment runs, and we are glad to review a draft imaging protocol against the outputs you need.