Voice Assistant for Outpatient Neurology Visits Based on Artificial Intelligence Technologies: A Multicenter Prospective Pilot Study
Voice Assistant for Outpatient Neurology Visits Based on Artificial Intelligence Technologies: A Multicenter Prospective Pilot Study
Documentation duties account for a substantial portion of an outpatient physician's working time and reduce time available for direct patient interaction. Voice assistants based on automatic speech recognition and large language models are being developed to automate medical documentation across clinical specialties. However, such ambient AI-based services have not been systematically validated in Russian-language outpatient neurology practice integrated with a regional electronic health record platform.
This pilot, multicenter, prospective, before-after study evaluated the feasibility and preliminary effectiveness of a voice assistant Service designed to automatically pre-fill the structured outpatient neurology visit protocol in the Moscow regional medical information system (EMIAS). The Service implements a pipeline of streaming speech-to-text transcription, two-speaker diarization, and large language model-based mapping of the dialogue between the physician and the patient onto the fields of the standardized neurology examination protocol.
Five neurologists at five outpatient clinics in Moscow participated. The study comprised three stages: (1) baseline timing of consultations without the Service; (2) timing of consultations with the Service after a two-week adaptation period, with parallel evaluation of transcription and pre-fill quality and of physician and patient satisfaction; and (3) statistical analysis. Three hundred twenty consultations were timed (160 per stage). A stratified random sample of 30 audio-recording / generated-protocol pairs was used to evaluate Service quality; free-text fields were rated on a 5-domain Likert questionnaire and on a 10-point visual analogue scale, and binary fields were rated dichotomously to derive sensitivity, specificity, accuracy, Jaccard index, and false-positive rate. Patient satisfaction was assessed by the modified Patient Satisfaction Questionnaire 8 (PSQ-8); physician feedback was assessed by a custom questionnaire (including the Net Promoter Score) and by semi-structured in-depth interviews with thematic analysis using grounded theory.
The primary outcomes were the change between stages in (a) the time of focused physician attention to the patient and (b) the time spent filling and editing the protocol. Secondary outcomes addressed total consultation time, transcription quality (Word Error Rate), expert-rated quality of pre-filled fields, patient satisfaction, and physician satisfaction.
The study was conducted under the framework of the Moscow Healthcare Department experiment on the use of digital innovation technologies in health care (Order No. 153 of 21 February 2025), and was approved by the local independent ethics committee.
Rationale and evidence gap. Documentation burden is widely recognized as a leading driver of professional burnout among physicians and reduces the time available for direct patient interaction. International evidence on ambient AI scribes - systems that combine automatic speech recognition (ASR) and large language models (LLM) - shows consistent reductions in documentation time and stable or improved patient satisfaction. However, evidence specific to Russian-language outpatient neurology practice integrated with a regional electronic health record is lacking. Neurologists are among the medical specialties with the highest reported burnout rates internationally; the specifics of Russian-language medical terminology, the structured neurological examination, and integration with a regional medical information system preclude direct extrapolation of international findings.
Technical implementation of the Service. The Service implements a distributed pipeline: streaming audio capture by an EMIAS client module; REST-API transmission of audio chunks to the speech recognition service with two-speaker diarization (physician/patient) and contextual re-clarification of streaming transcripts; asynchronous routing of transcription results through Apache Kafka topics to a pre-fill subsystem; large-language-model-based mapping of the full transcript onto a JSON schema of the structured outpatient neurology examination protocol (document class code 23951) using clinical reference dictionaries (including ICD-10 via the MKB10_ACTIVE_SHORT_INFO_V2 reference table); and return of the structured pre-filled protocol to the physician's user interface for review, editing, and signing. Audio is recorded via an active HD-capsule microphone placed at the physician's workstation (OGG Vorbis, 16 kHz, ≥ 64 kbit/s, mono). A priori quality targets fixed in the technical documentation before the pilot were: Word Error Rate of streaming transcription ≤ 10 %; sensitivity, specificity, accuracy, and Jaccard index for binary protocol fields ≥ 0.9; mean expert rating on a Visual Analogue Scale ≥ 8 out of 10; share of consultations with protocol generation time ≤ 35 seconds ≥ 90 %; physician Net Promoter Score ≥ 0 %.
Methodological design choices. A non-randomized single-group before-after design was selected as appropriate for a first-line feasibility evaluation in real clinical practice: the same physicians and clinics provide their own baseline reference in Stage 1, eliminating between-physician confounding while preserving ecological validity. A two-week post-deployment adaptation period was prespecified to mitigate purely learning-curve effects on Stage 2 metrics. Timing of consultation stages was performed by trained experts of the Scientific and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Healthcare Department on the basis of three synchronized data sources for each consultation: (i) automatic EMIAS event logs from protocol creation to signing; (ii) screen recording of the physician's workstation (VocoScreen NG, version 4.01.1); and (iii) webcam recording of the physician's face and gaze direction. The tri-source synchronization reduces systematic bias common to single-source or self-reported timing studies.
Quality evaluation framework. A stratified random sample of 30 paired audio recordings and Service-generated protocols (six per physician) was used to evaluate transcription and pre-fill quality. Reference (gold-standard) transcripts were prepared manually by an expert with more than five years of experience. Free-text fields of the protocol were rated by four independent neurology experts on a five-domain Likert questionnaire (Relevance, Accuracy, Completeness, Conciseness, Linguistic correctness) and on a 10-point Visual Analogue Scale; binary fields were rated dichotomously to construct a confusion matrix and derive sensitivity, specificity, accuracy, the Jaccard index, and the false-positive (hallucination) rate. Inter-rater agreement among the four experts was quantified by Gwet's AC2, preferred over Cohen's κ for categorical classifications with asymmetric prevalence. Word Error Rate was computed in Python with the jiwer library against the reference transcripts.
Analytical approach. Quantitative between-stage comparisons used Student's t-test or the Mann-Whitney U test depending on distribution, with Pearson's χ² for categorical variables and Cohen's d as the effect-size measure (α = 0.05, two-sided). Qualitative analysis of semi-structured in-depth interviews with each participating neurologist followed the grounded theory approach of Strauss and Corbin, with open, axial, and selective coding.
Inclusion Criteria:
Exclusion Criteria: