The CANVAS (Consensus And Norms for Medical Visualization Academic Standards) Study: Developing Professional Standards for Effective Medical Visualization - A Modified Delphi Study
The CANVAS (Consensus And Norms for Medical Visualization Academic Standards) Study: Developing Professional Standards for Effective Medical Visualization - A Modified Delphi Study
The goal of this observational study (modified Delphi consensus study) is to establish an international expert consensus on professional standards for medical visualization in medical visualization experts from from around the world, each with a minimum of five years of relevant professional experience. The specific objectives are:
Participants will:
Fill out a short background survey about their job title, training, work setting, location, and experience with AI image tools.
Rate 10 statements about medical visualization standards on a scale from 1 (Strongly Disagree) to 5 (Strongly Agree). These statements cover three topics: accuracy, accessibility, and ethics. Participants will complete up to 3 rounds of anonymous online surveys.
Write comments to explain why they disagree with any statement they rate 1 or 2.
Between rounds, review a summary showing their own earlier ratings and the group's overall results (with no names attached). Then, re-rate any statements the group has not yet agreed on.
This study uses a modified Delphi design to develop consensus-based professional standards for medical visualization. The standard Delphi method typically begins with open-ended questions to generate initial statements. In this modified version, a systematic literature review and a focus group replace that first open-ended round, reducing participant burden and producing a more structured starting questionnaire. The study follows the Conducting and Reporting of Delphi Studies (CREDES) framework and reports results according to the Accurate Consensus Reporting Document (ACCORD) checklist. All activities are conducted online. There is no physical study site.
Questionnaire Development The initial questionnaire was developed in two stages. In the first stage, a systematic literature review identified existing frameworks, guidelines, and evidence related to professional standards in medical visualization. Three databases were searched: PubMed (title and abstract fields), Embase (title field), and the Cochrane Library (title field). The search terms were "medical illustration" OR "medical visualization," limited to clinical studies and review articles published between May 2016 and May 2026. The research team also hand-searched professional guidelines from the Association of Medical Illustrators (AMI) and the Institute of Medical Illustrators (IMI), as well as all issues of The Journal of Biocommunication (JBC). A total of 20 articles were included for synthesis.
In the second stage, the principal investigator convened a focus group of domain experts, including clinicians and three practicing medical illustrators. The group held three online meetings to discuss the literature review findings and draft the initial statements. To protect the independence of the Delphi process, no focus group member serves on the Delphi expert panel. This process produced a 10-item questionnaire organized into three domains: (i) Accuracy and Scientific Rigor (4 items), (ii) Accessibility and Diversity (3 items), and (iii) Disclosure and Ethics (3 items).
Content Validity Assessment Before distribution to the Delphi panel, 10 cross-disciplinary experts rated the relevance of each item on a 4-point scale (1 = not relevant; 4 = highly relevant). Items with an item-level Content Validity Index (I-CVI) of 0.83 or higher were retained. Items scoring between 0.70 and 0.83 were revised. Items below 0.70 were deleted. None of these 10 validators serve on the Delphi expert panel.
Delphi Round Procedures The study includes up to three sequential rounds of anonymous rating, all conducted via the REDCap platform deployed by the Information Technology Office of National Taiwan University Hospital.
In Round 1, all panel members rate the 10 statements on a 5-point Likert scale (1 = Strongly Disagree to 5 = Strongly Agree). For any item rated 1 or 2, respondents provide an open-ended comment explaining their reasoning. Between rounds, the research team compiles descriptive statistics for each item. Each panel member receives (a) their own prior ratings, (b) anonymized group statistics (mean, median, interquartile range, and rating distribution), and (c) a summary of anonymized qualitative comments. Two research team members independently review qualitative comments using content analysis and draft revision recommendations. Items reaching consensus are finalized and removed from subsequent rounds. In Rounds 2 and 3, panel members re-rate only those items that have not yet reached consensus, in light of group feedback. No real-time interaction or group discussion takes place at any stage.
Consensus Definition An item reaches inclusion consensus when all three of the following criteria are met simultaneously: (1) 75% or more of panel members rate the item 4 or 5; (2) the group median is 4 or higher; and (3) the interquartile range (IQR) is 1 or less. An item reaches exclusion consensus if 75% or more of panel members rate it 1 or 2. Items that do not meet either threshold after three rounds are classified as non-consensus and reported descriptively.
Stopping Criteria The consultation process ends after a maximum of three rounds, regardless of item-level consensus status. If all items reach consensus before the third round, the process ends early. If the number of continuously participating members falls below 15 in any round, that round is considered unable to guarantee representativeness, and the Delphi process is terminated early.
Statistical Analysis Plan Descriptive statistics summarize expert ratings for each item, reported as mean ± standard deviation (SD), supplemented by median and IQR. All tests are two-tailed, with p < 0.05 considered statistically significant. Inter-rater agreement is assessed using Kendall's coefficient of concordance (W), interpreted as weak (W < 0.3), moderate (0.3 ≤ W ≤ 0.7), or strong (W > 0.7). All analyses are performed in IBM SPSS Statistics; graphical visualizations are produced in GraphPad Prism.
Experts were invited from medical illustration training programs or professional associations worldwide. Inclusion criteria required that invitees meet at least one of the following conditions:
Inclusion Criteria:
Exclusion Criteria:
Individuals meeting any of the above criteria were excluded to preserve the independence of the consensus process.
Taipei, Taipei 100, Taiwan