Medical Device Testing
Describes version 4.56.0
A purpose-built layer for human factors validation work on medical devices - critical tasks, participant demographics, use errors and close calls, and a draft HFE/UE report in the FDA submission outline.
Needs the meddevice entitlement (shown as Medical Device Testing in the Activation window). Without it, none of the surfaces in this guide appear and every study is a general research project.
Two things must both be true before these features show: your license includes Medical Device Testing, and the study itself is marked as one. A general study on a licensed machine looks exactly like it always did. Nothing in this guide leaks into ordinary research work, and recording behaves identically either way.
Marking a study as medical device work
- When you create a study, choose Medical device study / regulatory submission (FDA human factors guidance / IEC 62366-1) as the study purpose. Your last choice becomes the default for next time. The purpose covers the whole series - formative rounds and the summative alike; which one this study is, you declare with the study type in Study Setup.
- Already-created study? Change it under Study Setup → Study purpose. Switching a study off hides these features but deletes nothing. Everything recorded is still in the file if you switch back.
The question is about the deliverable, not your industry: pick it when the study supports a regulatory submission.
Critical tasks
A task is critical when its failure could cause harm. In Study Setup → Tasks, each task gains a criticality choice and, when critical, a box for what harm could result, that rationale goes straight into the report's critical-task section.
A task nobody has classified is reported as not yet assessed - deliberately distinct from not critical. The report will tell you how many tasks still need the judgement, rather than quietly treating an unexamined task as safe.
Participant demographics
A validation reports its results per user group, and often by age band, prior experience, or dexterity too. Define those characteristics and record each participant's values in Study Setup → Participants:
- Each characteristic is a row: its name, a picker for this participant's value, and an edit button that opens the value editor (add, rename, reorder, remove, with a warning that says how many participants a removal affects).
- Values are a fixed list rather than free text, so two participants who are the same are recorded the same way. That is what makes a per-group breakdown computable.
- Renaming or reordering values never disturbs who is already recorded; participants are stored against the value itself, not its wording.
- A participant left blank is reported as not recorded, never silently dropped from the counts.
Use errors, close calls and difficulties
Three kinds of thing can go wrong on a task attempt, and the report counts each: a use error (did something wrong, or failed to do something needed), a close call (caught it themselves and recovered), and a use difficulty (managed it, but with observable trouble).
There are two ways to record one, and they end in the same place:
- During the session, press Use error. While a task is running, the button under the task row opens the list of use errors this study already knows about. Choose the type, then press the line that matches. One press, no typing, and the moment is stamped from the recording clock. If it is not on the list, describe it under New discovered use error and it joins the list for every session after this one. The button waits for a running task, because a use error belongs to an attempt: anything that goes wrong between tasks is a note.
- Or in Review, from a note or from the Logging controls. Right-click a note in the Session data list and choose Create use error from this note…: pick which use error it is, how it counted, and add the root cause (usually what the participant told you in the debrief). The note stays a note; the use error is created from it and cites it. Or, while logging from the video, the Logging controls window has the same Use error… button the session has: while you are timing a task it files against that attempt when you stop, and with the playhead inside an attempt you already timed it files at once. Either way the use error belongs to a task attempt, so a note taken while no task was being timed asks you to time the task first.
Pressing during the session is the one to reach for when you already know what you are looking at, which after the first session or two is most of the time. Review is for everything you only understood afterwards. Neither is a lesser route: both write the same row, and either can be corrected in the catalogue.
Use errors show in Review's Session data list as their own kind of row, in red, at the moment they happened, with a Use errors filter chip once a study has any. Click one to scrub the video there. They are not edited from the list: open Analysis → Use Errors and use Edit… on the occurrence to change how it counted, what was written, or the root cause. They are in the study's data export too, as rows of type Use error.
A press during the session is written into the log stream as it happens, and filed against the attempt when the recording ends. Until then the participant counts in the list are of earlier participants, which is what makes them useful for recognising the right line.
A close call is a completed task, so it counts as a success in the outcome figures, and would be invisible there. FDA guidance treats close calls as evidence of thorough testing, which is why they are recorded and counted separately.
The same use error, made by two people
A validation report is read one use error at a time: what happened, how many participants did it, whether the URRA saw it coming, and what the harm would have been. All of that depends on one decision, which is whether the thing P02 did and the thing P05 did are the same thing.
Both routes ask it: Use error during the session, and Which use error is this in Review. The first time, describe it once; from then on it is in the list and later sessions join it with one press. If the study has a use-related risk analysis, its use errors are already in the list before anybody has done anything, so picking one of its lines is what links what you observed to what the URRA predicted. A use error with no line is a discovered use error: one the URRA did not have, which is the row a reviewer reads first. The list says discovered beside those and nothing beside the rest.
The list is ordered the way your URRA is, device-level lines first and then each task in order, with discovered use errors after their task's lines, and each line says how many participants are already in it. The count is there because it is the best cue that you have the right line; the order deliberately does not follow it, so nothing steers you toward the bin that is already filling.
Two things stay with the occurrence rather than the shared use error. The type is per person: P01 catching a ten-fold entry before confirming and P02 not catching it is one use error with two outcomes, which is what a close call is for. And the root cause is left empty for you to ask this participant, because the last one's reason is a claim about a different person, and root causes are what design changes are argued from.
The use-error catalogue
Analysis → Use Errors is the study's list: each use error with whether it is a discovered one, how many occurrences there are, and how many participants. Opening one shows the URRA line it matches as a table of every field, with not specified where the line has nothing written yet, and lists the occurrences behind that number, with who, when, how it counted for that person, what was written, and the cause they gave.
Each occurrence is a header row (who, which task, when, and a play button that takes Review to that point in that participant's recording) over its labelled attributes: type, what happened, and root cause, which reads not specified until you enter one. That is the whole argument for counting this way: a number you cannot watch is a number nobody can check. Occurrences pressed during the session, or promoted from a note, know the moment exactly. One typed into Review with no note behind it has no moment to know, so the play button offers the start of the attempt instead and its tooltip says that is what it is doing rather than implying a precision that was never recorded.
That last part is the point of the whole arrangement. Four people can make the same entry for three different reasons, and each reason points at a different design change. Grouping by cause would hide that; grouping by what happened shows it.
The picker reduces drift but cannot abolish it, so this is where a wrong press is undone. Checking means two different things, and which list you are in decides it. Checking two or more use errors in the left-hand list merges them, keeping the first name and repointing every occurrence. Opening one and checking its occurrences moves them: Move to… files them under a use error that already exists, Split off as a new use error… files them under one you name on the spot, and Unfile… takes the classification off without deleting anything, after asking. One occurrence or ten, it is the same gesture, so undoing a session's worth of wrong presses is one action rather than one per row.
A use error split off starts as discovered. It is a different error from the one those occurrences were filed under, so it does not inherit that one's risk-analysis line: link it afterwards if your analysis has a line for it. Splitting off every occurrence leaves the original in the catalogue with none, which the report prints as a use error nobody made; delete it unless the risk analysis has it.
Merge with URRA line… is for a discovered use error that turns out to be one the URRA already had: the press went to the wrong bin. Pick the line, and the dialog says what will happen in the same words the Merge dialog uses. If the line already has a use error, the discovered one merges into it: that one keeps its name, the occurrences move to it, and the report has one row for the line rather than two splitting a participant count. If the line has no use error yet, which happens when a line is added to the URRA after the list was seeded, the discovered one becomes its use error and keeps its occurrences; the button then reads Link. Either way it stops being discovered, because the URRA did have it. A use error that is a URRA line offers Split off every occurrence… instead, for when what was logged under it turns out to be something new: it checks every occurrence and opens Split, so they move to a discovered use error you name, while the predicted one stays in the catalogue as the line's use error, never observed. It is disabled when there are no occurrences. Lines the study itself added are not offered here; the use error each was made from is its use error.
Renaming is always safe: the recorded occurrences keep pointing at the use error, and only its name changes, so a better name reaches every table, citation and the report at once. A use error that matches a URRA line is that line, so renaming it in either place renames it in both. Deleting one that still has occurrences is refused, because the classification would be lost rather than moved.
Adding a discovered use error to the URRA
A discovered use error is the study's most consequential kind of finding, and often the URRA should have had it. Opening one in the catalogue offers Add to URRA: one press makes the URRA line from the use error's task and wording, links the two, and opens the line in Study Setup → URRA for you to write the potential hazardous situation, the potential harm and the risk control, which are yours to write and are empty until you do.
The line remembers where it came from. In the URRA editor it says it was added from this study's findings, and the report keeps both facts: section 6 shows the line in the table, marked, and lists the lines the study added; section 8 goes on reporting the use error as discovered, with the harm you wrote on the line, because for this validation it was discovered. Adding it does not make the URRA look as though it saw the error coming, which is what a reviewer would otherwise be misled by.
Ovo does not write into your risk file: the URRA in the study is yours to author and update, and a line the study added is a prompt to update the file, not a substitute for doing so.
Use-related risk analysis (URRA)
Study Setup → URRA records the use-related risk analysis (URRA) section 6 of the HFE report is built from: for each task, the potential use error, the potential hazardous situation and the potential harm as separate columns (the 2026 guidance's format), plus the risk control in place. A line can also be device-level - storage, packaging, rather than tied to one task. Each line can carry a reference into the manufacturer's risk file (an ID or clause); the risk file itself stays with the manufacturer - the study holds the linkage and the evidence, which mirrors how FDA reads submissions: the applicant authors the URRA, and observed use issues are reviewed against it.
That reading is built in: the report's section 6 shows the URRA table and then reads the validation's observed use errors against it - a use issue observed on a task the URRA doesn't cover is called out (it's the gap reviewers cite first), as is a critical task with no line. The format check raises the same findings while you work.
What a regulated study doesn't offer
Marking a study as medical device work takes three things away as well as adding several. This isn't a licensing limit - they're all still there in a general study - it's that each of them answers a question a submission doesn't ask, and would put the wrong kind of material next to the evidence.
- The Study Designer. It drafts a task set from a plain-language brief. A submission's tasks have to trace to the use-related risk analysis, so a protocol written from a description has the wrong provenance however well it reads. Build the tasks from the URRA instead.
- A/B designs. Comparing two designs against each other is a general research method. A regulated study validates one design against identified use-related risks, so the design picker and the per-design URLs would only invite a study that answers a different question.
- The AI findings report. It writes deliberately non-technical stakeholder prose and avoids severity and heuristic vocabulary - the opposite of what section 5 needs. The HFE/UE report is the artifact here, and two reports disagreeing about which one is the record helps nobody.
Everything else stays, surveys included: SUS and the after-task scales are ordinary instruments and a regulated study may still want them.
Design changes and the study series
These two are where a programme of rounds becomes traceable, which is why they belong to regulated work rather than to a single study. They are not offered in a general study: recording what a finding led to, and naming the later round that verified it, is section 5’s evidence rather than a note to self.
Design changes & recommendations
Every study ends in recommendations. Record them where the evidence lives. Analysis → Design Changes keeps a list of what the findings led to: each change or recommendation with a short description, a status as it moves from recommended to implemented to verified, and citations to the evidence that motivated it - the observer notes, participant quotes, and use errors it answers, picked from the study with a filterable chooser.
- Citations resolve live: a cited note always shows what it currently says, and if a cited session is deleted the citation goes with it while the change itself survives.
- Ending at recommended is a perfectly good final state - a study that hands its recommendations onward is finished.
- In a regulated study, recorded design changes also render in the HFE/UE report's section 5 with their status and cited evidence.
Study series: several studies, one line of work
Most products aren't tested once. You run a round, change something because of what you saw, run another round to find out whether the change worked, and eventually run a validation. Each of those is its own study file, and a study series is the thing that holds them together.
Open Study Series on the Analysis page. Create a series (it's a single file, so put it wherever you keep the studies), then Add study… for each study it covers. Every study you add becomes a round, in the order you arrange them.
- Give each round a name and a purpose: "Round 1", "First look at the instructions for use". These are your words, and re-reading a study never overwrites them.
- The facts come from the study: how many participants and in which user groups, how many sessions and tasks, how many use errors. Ovo Logger reads them when you add the study and shows you when it last did. Click Re-read the study after you've done more work in it.
- Chain the design changes. Each round lists the design changes recorded in that study, and against each one you can say which later round verified it. Only rounds that come after are offered - a round that ran first cannot have verified a change made later.
- See how each round updated the risk analysis. Each round also lists the lines its findings added to the URRA: the discovered use errors you added from that study's catalogue (see adding a discovered use error). They are read from the study along with everything else, so re-read the study after adding one.
That chain - found here, changed, verified there - is what a series of formative rounds amounts to, and it's the one thing no single study file can record about itself.
The round stays, reporting what was read and saying when it was read. Findings are meant to outlive the raw material: research data gets purged on a retention schedule long before the conclusions drawn from it stop mattering. You just can't re-read a study that isn't there any more.
For medical-device work, the series is what fills in section 5 of the HFE/UE report - the preliminary analyses and evaluations. Add the validation study to the series and section 5 writes itself: the round table, the design-change chain, and how the evaluations updated the URRA. See the medical-device guide.
The HFE/UE report
Analysis → HFE/UE Report opens the report as a workspace: the draft renders in the eight-section outline from FDA's Content of Human Factors Information in Medical Device Marketing Submissions, built fresh from the study's current data, with the narrative sections editable in place.
- Click Write this section… (or Edit…) on a narrative section and it becomes editable right in the document - bold, italic, lists, headings and quotes, with guidance above it saying what the section is expected to contain. Save lands it in place.
- Photographs and diagrams go in the document. Section 3 is expected to show the device's controls, displays and labelling, so the editing toolbar has an image button: pick a picture, describe it for screen-reader users, add a caption if you want a figure number, and it is placed at your cursor. The image is stored inside the study and embedded in the exported file, so the report stays a single document you can mail. See the main guide.
- Only the sections the study cannot generate are editable; every figure and table is computed and cannot be changed here or anywhere.
- Export HTML… writes the clean document beside the study and opens it in your browser, no editing controls, no workspace markers, ready to send. The file is laid out one element per line and indented, so it reads in an editor as well as in a browser.
- Two sources write the report, and every heading says which: Ovo for what the software produces from the study data, HFPro for what the human factors professional writes (the device description, known use problems and the conclusion are always HFPro), and Ovo, not yet when a section needs something the study doesn't hold yet - the line under the heading says what. A "who writes what" table at the top is built from the same marks. All of this is scaffolding for the author, shown on an amber ground in the window; the exported report carries none of it, just the headings, the content, and one line saying it is a draft that requires the HFPro's review and sign-off.
- Generated sections include the participant demographics breakdowns, the critical-task table with rationales, task outcomes, and use errors counted per task.
- One row per named use error, headed by how many participants it happened to rather than how many times it happened - one person doing the same thing twice is one person the device failed, and both figures are shown. Each row names the risk-analysis line it matches, or says discovered, which is the row a reviewer reads first; the potential harm comes across from the line. A discovered use error you have since added to the URRA stays discovered here, marked as added, because for this validation it was. Use errors your analysis has that nobody made are listed too, at zero, because a predicted error that never happened is a result. Where one use error was recorded with several different root causes, section 8 says so: each reason points at a different design change.
- Every use issue is also accounted for individually, critical tasks first: what happened, which participant and user group, how the attempt was scored, the root cause, whether your risk analysis predicted that task's failures - called out loudly when a critical task's issues were not predicted, and which design change answers it. Nothing there is new information; it is everything you have already recorded, read together the way a reviewer asks about it. The judgement of whether the remaining risk is acceptable stays yours, in the conclusion.
- Results per user group. Mark one demographic characteristic as the user-group dimension (edit the characteristic in the Participants editor and check “This characteristic defines the user groups”) and section 8 reports its outcome tables per group - the breakdown a summative is required to show, with participants missing a group value collected into a visible “no user group recorded” bucket rather than silently pooled. The format check also counts group sizes against FDA reviewers’ stated expectation of at least 15 participants per distinct user group.
- Declare the validation. Study Setup carries a study type picker for regulated studies. Choose Summative validation for the submission's HF validation and the validation rules switch on: group sizes are checked against each group's recruitment plan (set per value on the user-group characteristic) and fall short as blocking findings, every critical task must be exercised by every user group, and the report opens section 8 with the protocol facts - the frozen device/UI version you state in Study Setup, each participant's training arm (trained, untrained by design, or not recorded - set in the Participants editor with the training date and materials), and the training-to-testing gap computed from those dates. Formative studies see none of this.
- A regulated study starts set up for this. Choosing the regulatory-submission purpose (at creation, or later in Study Setup) seeds a marked “User group” demographic ready for your group values, and labels four quicklog tags to match the reporting vocabulary: U Use error, C Close call, D Use difficulty, H Help given, so live observer notes land pre-aligned with what section 8 counts. Nothing you have already named or labelled is ever overwritten.
- Section 5 comes from the study series. The preliminary analyses and evaluations span several studies - the formative rounds that led to the design you validated, so they cannot come from one study file. Group the studies into a study series (Analysis → Study Series; see the main guide) and add the validation study to it: section 5 then generates the round table, when, who took part and in which user groups, tasks, use errors - followed by each round's design changes and, where you have said so, the later round that verified them, and then how the evaluations updated the URRA: the lines each round's findings added, which is what the guidance asks a sponsor to show. The validation itself is not listed there, since section 5 is the preliminary work and section 8 is where the validation is reported; it stays available as the round a change can be marked verified in. A round whose study file has since been deleted or archived still reports, from what was read, with the date it was read.
- The file is self-contained HTML - mail it, print it, open it anywhere.
Checking your writing against the required format
Check format reviews the report against the FDA outline and shows its findings in the document, beside the section each one is about. Two kinds of check run:
- Format check - computed. Exact checks Ovo Logger does itself: sections not yet written, whether section 2 mentions all four of users, uses, environments and training, numbers in the conclusion that match no figure the study computed, critical tasks with no scored result, use errors with no root cause, and occurrences not yet filed under a named use error. These run instantly and need no AI Pack.
- ✨ AI review - advisory. With the AI Pack installed, the same click also has the AI read each written section against that section's required content and suggest what is missing, too vague, or not supported by what the text says, including whether recorded root causes actually explain why an error happened rather than restating what happened. The AI never writes or rewrites anything; the failure mode of a wrong suggestion is a moment of your time, and a qualified person decides every change.
Checking again after you revise doesn't start from scratch: sections you haven't touched keep their previous review unchanged, and where you've addressed an earlier finding the new review says so - an “Addressed since the last review” note listing exactly which items your revision answered. The note quotes the earlier review word-for-word - the AI only points at which of its items were addressed; Ovo Logger supplies the wording and discards anything that doesn't correspond to a real earlier finding, so the credit is real. The AI review runs only when you click. Nothing reviews your writing while you're still authoring.
Findings never appear in the exported document. They belong to the workspace.
The report is a draft for a qualified human factors professional to complete, review and sign. Every number in it is computed by Ovo Logger from the study's own data, no AI writes any figure, but the document itself is a starting point, not a submission.