Open-access Natural Language Processing applied to electronic records: surveillance and detection of health events

Abstract

Text fields in medical records are a valuable source for Public Health Surveillance but remain underutilized. This study describes the use of natural language processing (NLP) to enhance the identification of suspected cases and monitor disease trends in electronic records from the Urgency and Emergency Network (Rede de Urgência e Emergência - RUE), in the municipality of Rio de Janeiro (MRJ). Texts were pre-processed, and rules were applied to identify individual (measles and rubella) and collective (diarrhea and influenza-like syndrome) events, comparing the results with ICD-10 data from January 2023 to September 2024. A total of 28 suspected measles cases and 33 suspected rubella cases were identified through ICD, while the NLP technique detected an additional 30 suspected cases of measles and 17 of rubella based on patient complaints. Time series of diarrhea and influenza-like syndrome (síndrome gripal - SG), stemming from ICD and complaints, showed a cross-correlation above 0.93 at lag 0. Complaint analysis, particularly after the discontinuation of nonspecific SG ICD codes by RUE management, revealed a greater stability and expanded detection of suspected cases, demonstrating the potential of NLP in epidemiological surveillance in MRJ.

Keywords:
Natural Language Processing; Public Health Surveillance; Electronic Health Records; Epidemiological Monitoring

Resumo

Campos textuais de prontuários são fontes ricas para a Vigilância em Saúde, mas ainda pouco exploradas. Este estudo descreve o uso de processamento de linguagem natural (PLN) para ampliar a identificação de casos suspeitos e monitorar tendências de doenças em registros eletrônicos da Rede de Urgência e Emergência (RUE), no município do Rio de Janeiro (MRJ). Os textos foram pré-processados e aplicou-se regras para identificar eventos individuais (sarampo e rubéola) e coletivos (diarreia e síndrome gripal), comparando os resultados com dados da CID-10 entre janeiro de 2023 e setembro de 2024. Identificou-se 28 casos suspeitos de sarampo e 33 de rubéola pela CID, enquanto a técnica de PLN detectou mais 30 casos suspeitos de sarampo e 17 de rubéola, a partir das queixas dos pacientes. Séries temporais de diarreia e síndrome gripal (SG) construídas com CID e queixas mostraram correlação cruzada acima de 0,93 em lag 0. A análise das queixas, especialmente após a descontinuidade de CIDs inespecíficos de SG pela gestão da RUE, revelou maior estabilidade e ampliação na detecção de casos suspeitos, evidenciando o potencial do PLN na vigilância epidemiológica do MRJ.

Palavras-chave:
Processamento de Linguagem Natural; Vigilância em Saúde; Registros Eletrônicos de Saúde; Monitoramento Epidemiológico

Resumen

Los campos textuales de registros médicos son valiosos para la Vigilancia en Salud Pública, pero poco explotados. Este estudio analiza el uso del procesamiento de lenguaje natural (PLN) para identificar casos sospechosos y monitorear tendencias de enfermedades en registros electrónicos de la Red de Urgencias y Emergencias (RUE) en Rio de Janeiro (MRJ). Los textos fueron preprocesados y se aplicaron reglas para identificar eventos individuales (sarampión y rubéola) y colectivos (diarrea y síndrome gripal), comparando resultados con datos de la CIE-10 entre enero de 2023 y septiembre de 2024. Se identificaron 28 casos sospechosos de sarampión y 33 de rubéola mediante la CIE, mientras que el PLN detectó 30 casos sospechosos adicionales de sarampión y 17 de rubéola a partir de quejas. Las series temporales de diarrea y síndrome gripal (SG) basadas en CIE y quejas mostraron correlación superior a 0,93 en lag 0. El análisis de quejas, tras la eliminación de códigos CIE inespecíficos de SG por la RUE, mostró mayor estabilidad y amplió la detección de casos, destacando el potencial del PLN en la vigilancia epidemiológica en el MRJ.

Palabras clave:
Procesamiento de Lenguaje Natural; Vigilancia en Salud; Registros Electrónicos de Salud; Monitoreo Epidemiológico

Introduction

Electronic health records (EHR) are systematized collections of data containing signs and symptoms, test requests and results, images, prescriptions, and other individual patient information collected by healthcare professionals. The framework of these records includes fields in a structured format, such as numeric fields and check boxes, and unstructured fields, such as free-text boxes. The analysis of these textual fields, using Natural Language Processing (NLP) techniques, can incorporate new insights into health surveillance performance and has been useful in supporting decision-making and benefiting health at individual and collective levels1,2.

Exploratory approaches applied to free-text fields through NLP have contributed to the understanding of semantic relationships in medical texts, with the possibility of extracting information, segmenting sentences, and categorizing and associating words3. Previous studies using NLP have, for example, enabled the identification of comorbidities in diabetic patients, the monitoring of suicide attempts, the detection of influenza cases, and the identification of emerging syndromes in the pre-syndromic surveillance approach4-7.

However, exploring free-text fields in an automated and reliable way requires mastering computer language text interpretation techniques. NLP originated from the mixture of linguistics and computing8 and currently refers to the ability to translate human language to generate new insights, through automated analysis, involving statistical computational or machine learning techniques9. Basically, different processes for extracting and detecting information in free-text fields involve stages of collecting, processing, analyzing, and interpreting information. In this sense, NLP approaches are diverse and can use everything from traditional classification models, such as systems based on rules using regular expressions or keyword searches, to advancing in the development of more sophisticated machine learning (ML), deep learning (DL), and Large Language Models (LLM) algorithms10,11.

In the Municipality of Rio de Janeiro (MRJ), the use of emergency and urgent network EHRs began during the COVID-19 pandemic. Currently, these records support disease detection activities and the monitoring of temporal trends of important illnesses (such as arboviruses and influenza syndrome), with the identification of events and epidemiological monitoring consolidated in a warning panel. These activities were initially proposed based on the analysis of codes from the Tenth Revision of the International Classification of Diseases (ICD-10), contained in structured fields of these electronic records12.

Despite the advances obtained with the analysis of ICD-10 codes for health surveillance (HS), the content of the free-text field, which allows one to search for signs, symptoms, and possible diagnoses in the patient’s complaint record, remained unexplored. Considering the need to diversify the processes adopted for early detection and reinforce the monitoring of relevant public health events, this article aimed to describe strategies to expand the identification of suspected cases and strengthen the monitoring of trends in diseases of interest in public health, through the use of natural language processing, applied to EHR, in the MRJ.

Methodology

Data source

The present study is based on electronic records of care provided at 15 Emergency Care Units and 4 Regional Emergency Coordination Units, between January 2023 and September 2024. These units represent 82.0% of the services that make up the emergency care network in Rio de Janeiro’s municipal sphere.

Sociodemographic variables, health unit data, identification of care according to ICD-10, and free-text fields for recording patient complaints were analyzed. The records were captured daily, in an automated manner, using connections via the Application Programming Interface (API). The records were uploaded to a local database at the Epidemiological Intelligence Center (Centro de Inteligência Epidemiológica - CIE)13 .

All 4,089,949 available for the analysis period were included, as were the fields of complaint and ICD-10 achieved completeness rates of 91.5% and 99.8%, respectively. The average number of monthly visits per unit was 10,250 (standard deviation of 2,611), indicating some variability between units. This variability reflects the difference in demand and capacity of services, as well as the population residing in the areas where each unit operates. Appointments for which the complaint field was not completed were treated as missing data.

Preprocessing

A sequence of modifications was made in the patient complaint field to improve the capability of analyzing the texts. Procedures, such as removing numbers and special characters, converting to lowercase, and removing stopwords (non-relevant terms such as “the”, “in”, “to”, etc.), were applied. These techniques are widely used in rule-based classification methods, known as keyword-based, rule-based, or linguistic systems. They are typically based on a manually constructed keyword dictionary, together with a series of syndrome mapping rules and a preprocessing module to normalize lexical variants14.

The symptoms for creating rules with regular expressions to capture cases from complaints were defined by consensus by a multidisciplinary team from the CIE, which consisted of two physicians, a nurse, and three epidemiologists. The definition of a suspected case of events of interest, based on the Notifiable Diseases Information System (Sistema de Informação de Agravos de Notificação - SINAN) Notification Form, and the language and terms used in routine patient care were considered.

The list of stopwords was based on the tidytext v0.4.2 package15, a stopwords-iso source, and modified from the sum of the high-frequency words identified in the database, with no meaning added to the complaints. The process of standardizing the terms considered possible grammatical variations identified in the database itself so that, for example, {“febrile”, “high temperature”, “feverish”, “pyretic”, “hyperthermia”, “hot body”, “high temp”, “fev”} could be translated as “fever”. A boolean variable was created for each symptom, obtaining a encounters classification concerning the symptoms listed using regular expressions. Chart 1 presents the list of symptoms and their grammatical variations in writing, disregarding accents and capital letters, selected in the elaboration of regular expressions for the standardization of symptoms in the complaints field.

Chart 1
Symptoms, possible variations, and synonyms used to standardize terms in preprocessing records.

Encounters classification

After the fields were prepared and standardized, the encounters were classified according to the search for regular expressions based on the rules of each disease in the two categories of interest: individual care events (measles and rubella) and collective care events (diarrhea and influenza syndrome). The choice of attention events was aligned with criteria of impact on public health, potential for dissemination, and vulnerability16. Both categories were defined based on distinct actions aimed at managing individual suspicions and case clusters. The suspected diagnosis of diseases in the process of elimination, such as measles and rubella, is even more important than the confirmation of cases. Furthermore, these diseases are less frequent when compared to events requiring collective attention. In this category, nevertheless, the aim is to expand the signals to identify trends in diseases, such as diarrhea and influenza syndrome, which are mainly reported in the form of outbreaks in the territory, preventing the identification of isolated cases through SINAN.

Chart 2 presents the rules defined as logical operators “AND”, “OR”, and “NOT”, representing the regular expressions used, for the classification of encounters for diseases categorized as individual care events (measles and rubella) and collective care events (influenza syndrome and diarrhea).

Chart 2
Rules applied to electronic records for text search, by health event.

The terms used for the text search in Chart 2 included the set previously standardized and presented in Chart 1 and other specific terms that direct the classification of the condition of interest, such as “ganglion” and “lump”.

Particularly, individual attention events were captured and sent by email to the MRJ rapid response units (RRU) to validate the event and carry out possible HS action developments in the area.

The sequence of steps performed for preprocessing and classification of encounters based on patient complaints contained in the EHR is summarized in Figure 1.

Figure 1
Flowchart of processing stages applied to the complaint field in MRJ emergency and urgent care system electronic records.

Comparison between ICD classification and text fields

The number of suspected rubella and measles cases identified by the ICD-10 structured field was compared to the total number of events captured by the set of textual rules, and the percentage increase was presented.

The time series of services classified by ICD as diarrhea and influenza syndrome were compared to the series generated using NLP through cross-correlation using the differentiated series. This methodology involves calculating correlation coefficients for different time lags, in which one series is kept fixed and the other is shifted forward and backward in time. With lags ranging from -15 to +15 epidemiological weeks, the lag with the highest positive or negative correlation was identified, considering a 95% confidence interval (95%CI).

Data analyses were conducted in R language version 4.4.017 using the stats17, ggplot218, tidyverse19, duckdb20, sendmail21, and VennDiagram22 packages. The codes used in service preprocessing and classification are available in an open repository23.

Ethical considerations

The use of data from the urgency and emergency network was contemplated in a project approved under opinion no. 6572784 of the Research Ethics Committee of the Rio de Janeiro Municipal Health Department.

Results

Based on the searches carried out during the analysis period, the results of individual and collective care events are presented, both totaling the searches currently carried out using ICD and new searches using the NLP technique.

Figure 2 shows the cumulative total of suspected measles and rubella cases identified by the ICD category, complaints (textual rules), and both approaches, from January 2023 to September 2024.

Figure 2
Individual attention events captured by the ICD (structured field) and complaint (unstructured field) in the MRJ, from January 2023 to September 2024.

During the analyzed period, 28 suspected cases of measles and 33 of rubella were identified from the structured ICD field. After applying the textual rules to the complaints, 30 cases of measles and 17 of rubella were captured, with one case of measles and two cases of rubella being identified by both approaches. In total, 57 suspected cases of measles and 48 of rubella were identified, representing an increase of 103.5% and 45.4%, respectively, as compared to the isolated search by ICD.

The application of NLP focused on monitoring time series and trends in collective attention events was compared to the series prepared from specific ICDs, in which isolated suspected cases are no longer sought, but rather possible changes in the behavior of certain illnesses, such as diarrhea and influenza syndrome.

Figure 3 presents the time series of care services classified as diarrhea by both ICD and complaints, as well as the cross-correlation between the two time series, by epidemiological week (EW).

Figure 3
Time series for diarrhea using the ICD and complaint fields and cross-correlation between the series in the MRJ, from January 2023 to September 2024.

As observed in Figure 3a, the number of diarrhea cases, from the structured ICD field (dashed gray line), behaved similarly to the number of cases identified by the rules obtained from the patient complaint field (black line). The maximum value of the cross-correlation between the differentiated series at lag 0 equal to 0.93, at a 95%CI (Figure 3b), suggests correspondence between the series in time, without significant delay or anticipation when using both search approaches.

Between EW 07 and EW 12 (2024), an increase was observed in the series originating from the textual field (complaints) to those classified by the ICD, reaching a maximum of 6,630 weekly cases, a number higher than that recorded by ICD in EW 50 of 2023, of 5,822 cases. The opposite can be seen, from EW 17 of 2024, when cases identified by ICD were higher than those captured by the complaint field until the end of the time series.

Figure 4 presents the time series and cross-correlation between the services classified by both the ICD and complaint fields as influenza syndrome.

Figure 4
Time series for influenza syndrome using the ICD and complaint fields and cross-correlation between the series in the MRJ, from January 2023 to September 2024.

As highlighted for diarrhea monitoring, Figure 4 shows that the influenza syndrome series captured by the approaches with structured (ICD) and unstructured (complaints) data also exhibited similar patterns over time, with the cross-correlation reaching the maximum value of 0.95 at lag 0, at a 95%CI (Figure 4b). This means that they correspond in time, with no significant delay or advance between the series. Nonetheless, throughout the entire analyzed period, it is possible to observe that the number of cases captured by complaints is higher than that of the ICD approach, primarily until the beginning of EW 44 of 2023.

Two important milestones in the context of information recorded in the electronic medical records of the urgency and emergency network were highlighted in Figure 4a (dotted lines). The first, starting on EW 44 of 2023, indicates the discontinuation of the use of ICD R05 (acute or chronic cough), and the second, starting on EW 11 of 2024, the discontinuation of the use of ICD B34.9 (unspecified viral infection) as per guidance from the management of the MRJ urgency and emergency network, both nonspecific ICDs commonly used for the classification of influenza syndrome. Thus, the effect could be observed in the shortest distance between the series from the milestones indicated in the graph, still showing a predominance of cases captured by the complaint field.

Discussion

Traditional epidemiological surveillance systems were designed, for the most part, using a passive paradigm, focused on counting the occurrence of human cases, hospitalizations, positive laboratory diagnoses, pathogen genomes, and deaths, and are based mainly on structured data24. While these systems are useful for detecting diseases and monitoring event trends, they have significant limitations, such as the possibility of underreporting, delayed response times, and the inability to capture detailed information about emerging events, especially those that are not immediately recognized as public health threats25.

The diversification of data sources for the development of public policies, particularly for the surveillance of communicable diseases, has been encouraged by different authors26,27 and is already a reality for the work processes of the MRJ HS team. Structuring a health system with electronic medical records in its own healthcare network was decisive for the latest advances in public health in Rio de Janeiro, allowing access to data in “almost real” time and in an automated manner. Although the processing of the textual content of clinical care is a valuable resource, it was not considered for the incorporation of emergency and urgent care system resources and for the use of ICD codes for detecting and monitoring events by HS, processes that demand high sensitivity12,28.

The results of the present study indicated that the process of automatic classification of signs and symptoms of complaints reinforced and expanded the capacity for the timely capture of individual attention events. In fact, by incorporating patient complaints into the strategy for searching for suspected cases in the MRJ emergency and urgent care system, it is possible to confirm true progress, when compared to that detected by ICD-10, with the identification of approximately 100% and 45% more suspected cases of measles and rubella, respectively. Capturing these events enabled actions at an individual level, such as collecting clinical data and biological samples, sending them to the laboratory, identifying contacts, and implementing vaccinations16. The NLP also contributed to achieving the goals related to the quality indicators of epidemiological surveillance of exanthematous diseases, aiming to recover the measles elimination certificate in Brazil29.

In the context of collective attention events, it was found that the series created from the text field were strongly correlated with the time they originated from the ICD-10 codes. The conceptual hypothesis that processing the unstructured field in non-traditional sources for HS constitutes an effective approach for identifying changes in trends and reinforces the ICD-oriented strategy was confirmed. As shown, the influenza syndrome series generated by the unstructured field (complaints) remained stable when compared to the monitoring carried out using the structured field, which is directly susceptible to the code entered by the health professional in the medical record. Furthermore, changes in the use of ICD codes are common situations in the routine of emergency care, influenced by the epidemiological context and by management guidelines, and this can affect strategies that rely on completing this field.

The strategy adopted for the text field was also able to capture a greater volume of data in a large section of the analyzed series. The free texts in the electronic medical records, as suggested by the findings presented here, contained rich and detailed descriptions of patients’ symptoms, offering information that traditional models could rarely capture. The analyses carried out for diarrhea and influenza syndrome, aligned with monitoring by the ICD, reinforced the warning signal and resulted in local developments, such as the expansion of the collection of biological samples and the preparation of technical notes to the HS network to strengthen syndromic surveillance in the MRJ.

Although gains were observed for health protection activities, the obtained signals did not anticipate the series verified by the ICD, even if one considers that the complaints represent the conditions in a more diversified manner, as reported in other studies30,31. It is worth noting that seeking to anticipate signals and trends within the same set of data is a challenge.

The anticipation of signals in different data sources was evidenced when using a Bayesian classifier32, based on the probability of each word (signs and symptoms) for each syndrome, detecting three respiratory outbreaks with high sensitivity and specificity31. In this study, the time series of complaints in an emergency system showed a correlation with hospital admissions, advancing them by an average of 10.3 days. In the context of respiratory syndromes, a gain of two weeks was shown, based on the analysis of Primary Health Care (PHC), when compared to cases reported to surveillance systems. However, the strategy was based on structured data30. In this sense, incorporating complaints from PHC patients into the warning model proposed here could enhance the response capacity, whether in the stability of signs and symptoms captured by the healthcare system, or for the preparation of the hospital network, by previously pointing out changes in trends about hospitalization data in the MRJ.

The lack of records from state units or from those that do not have a connection to the application interface was a limitation identified with the use of data from the urgency and emergency network. Access to records was limited to units with an API connection available, which restricted the reach of the municipality’s epidemiological scenario analysis. However, representativeness was not strongly compromised, since the data volume corresponded to 82.0% of the units in the MRJ emergency and urgent care system, covering nine of the city’s ten planning areas. At this point, it should be reiterated that the proposed surveillance fulfills its role in terms of sensitivity to changes in temporal trends, in addition to favoring early detection and prevention of new problems, an essential interest of HS33. When considering the dynamics of communicable diseases (such as COVID-19) and the importance of identifying areas (focus/index) for interventions in the transmission chain, the adopted strategy allowed for the timely targeting of control actions.

As it is a new methodological approach in the practice of health surveillance at MRJ, the rule-based NLP method allowed for progress in the use of previously unexplored unstructured data and represented a new window of possibilities for the CIE in the use of alternative sources, from the perspective of syndromic surveillance and incorporation of digital innovations for the SUS. Given the advancement of analyses using more sophisticated methodologies in the literature, the current model may still not be able to fully capture contextual and regional variations in complaints, as observed in a recent study.6 The limitations identified in this study were characterized by the high demand for manual work, as it requires the creation and continuous adjustment of an extensive set of rules to achieve adequate generalization. Furthermore, developing specific rules is a time-consuming process and susceptible to constant updates, as the code requires frequent maintenance to keep up with changes in language usage, variations, and linguistic exceptions.

Conclusion

The study of the patient complaint text field showed greater stability in the temporal series analysis and greater scope in detecting suspected cases of measles and rubella in the MRJ concerning monitoring by the ICD structured field. As future perspectives, the further development of NLP techniques using ML, DL, and LLM models is already part of the CIE planning, to extract signs and symptoms in an even more sophisticated way and considering grammatical contexts not yet explored by the technique defined by rules.

The integration of data from other sources, such as PHC, laboratory, and hospital data, which are key pieces to improve the capacity to monitor trends and trigger early warnings from open fields, will be a major task. These advances promise to expand the effectiveness of surveillance by capturing local variations and outbreaks at early stages, facilitating more agile and targeted responses.

Acknowledgements

To Dr. Daniel Ricardo Soranz, Municipal Secretary of Health of the city of Rio de Janeiro. To the Pan American Health Organization (PAHO) and the Ministry of Health.

References

  • 1 Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, Kang J. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 2020; 36(4):1234-1240.
  • 2 Paul MM, Greene CM, Newton-Dame R, Thorpe LE, Perlman SE, McVeigh KH, Gourevitch MN. The State of Population Health Surveillance Using Electronic Health Records: A Narrative Review. Popul Health Manag 2015; 18(3):209-216.
  • 3 Xiao W, Jing L, Xu Y, Zheng S, Gan Y, Wen C. Different Data Mining Approaches Based Medical Text Data. J Healthc Eng 2021; 2021:1285167.
  • 4 Bey R, Cohen A, Trebossen V, Dura B, Geoffroy PA, Jean C, Landman B, Petit-Jean T, Chatellier G, Sallah K, Tannier X, Bourmaud A, Delorme R. Natural language processing of multi-hospital electronic health records for public health surveillance of suicidality. Npj Ment Health Res 2024; 3(1):6.
  • 5 Chen C, Zheng X, Liao S, Chen S, Liang M, Tang K, Yin M, Liu H, Ni J. The diabetes mellitus multimorbidity network in hospitalized patients over 50 years of age in China: data mining of medical records. BMC Public Health 2024; 24(1):1433.
  • 6 Nobles M, Lall R, Mathes RW, Neill DB. Presyndromic surveillance for improved detection of emerging public health threats. Sci Adv 2022; 8(44):eabm4920.
  • 7 Ferraro J, Ye Y, Gesteland P, Haug P, Tsui F, Cooper G, Van Bree R, Ginter T, Nowalk AJ, Wagner M. The effects of natural language processing on cross-institutional portability of influenza case detection for disease surveillance. Appl Clin Inform 2017; 8(2):560-580.
  • 8 Nadkarni PM, Ohno-Machado L, Chapman WW. Natural language processing: an introduction. J Am Med Inform Assoc JAMIA 2011; 18(5):544-551.
  • 9 Osman M, Cooper R, Sayer AA, Witham MD. The use of natural language processing for the identification of ageing syndromes including sarcopenia, frailty and falls in electronic healthcare records: a systematic review. Age Ageing 2024; 53(7):afae135.
  • 10 Harris J, Laurence T, Loman L, Grayson F, Nonnenmacher T, Long H, WalsGriffith L, Douglas A, Fountain H, Georgiou S, Hardstaff J, Hopkins K, Chi YL, Kuyumdzhieva G, Larkin L, Collins S, Mohammed H, Finnie T, Hounsome L, Riley S. Evaluating Large Language Models for Public Health Classification and Extraction Tasks. arXiv 2024 [Internet]; [cited 2024 set 30]. doi: https://doi.org/10.48550/arXiv.2405.14766
    » https://doi.org/10.48550/arXiv.2405.14766
  • 11 Sim J ah, Huang X, Horan MR, Stewart CM, Robison LL, Hudson MM, Baker JN, Huang IC. Natural language processing with machine learning methods to analyze unstructured patient-reported outcomes derived from electronic health records: A systematic review. Artif Intell Med 2023; 146:102701.
  • 12 Morais JHA, Cruz DMDOE, Saraceni V, Ferreira CD, Aguilar GMO, Cruz OG. O uso de fontes não-tradicionais para a vigilância em saúde: atendimentos de urgência para detecção precoce de eventos [Internet]. SciELO Preprints 2024; [cited 2024 set 30]. doi: https://doi.org/10.1590/SciELOPreprints.8996
    » https://doi.org/10.1590/SciELOPreprints.8996
  • 13 Cruz DO, Ferreira CD, Carvalho LF, Saraceni V, Durovni B, Cruz OG, Garcia MHO, Aguilar GMO. Inteligência epidemiológica, investimento em tecnologias da informação e as novas perspectivas para o uso de dados na vigilância em saúde. Cad Saude Publica 2024; 40(8):e00160523.
  • 14 Conway M, Dowling JN, Chapman WW. Using chief complaints for syndromic surveillance: A review of chief complaint based classifiers in North America. J Biomed Inform 2013; 46(4):734-743.
  • 15 Silge J, Robinson D. tidytext: Text Mining and Analysis Using Tidy Data Principles in R. J Open Source Softw 2016; 1(3):37.
  • 16 Brasil. Ministério da Saúde (MS). Guia de Vigilância Epidemiológica. Brasília: MS; 2009.
  • 17 R Core Team. R: A Language and Environment for Statistical Computing [Internet]. Vienna: R Foundation for Statistical Computing; 2024 [cited 2024 set 30]. Available from: https://www.R-project.org/.
  • 18 Wickham H. ggplot2: Elegant Graphics for Data Analysis. New York: Springer-Verlag; 2016.
  • 19 Wickham H, Averick M, Bryan J, Chang W, McGowan L, François R, Grolemund G, Hayes A, Henry L, Hester J, Kuhn M, Pedersen TL, Miller E, Bache SM, Müller K, Ooms J, Robinson D, Seidel DP, Spinu V, Takahashi K, Vaughan D, Wilke C, Woo K, Yutani H. Welcome to the Tidyverse. J Open Source Softw 2019; 4(43):1686.
  • 20 Mühleisen HR. duckdb: DBI Package for the DuckDB Database Management System. 2024.
  • 21 Mersmann O. sendmailR: Send Email Using R [Internet]. 2009 [cited 2024 out 15]. p. 1.4-0. Available from: https://CRAN.R-project.org/package=sendmailR
    » https://CRAN.R-project.org/package=sendmailR
  • 22 Chen H, Boutros PC. VennDiagram: a package for the generation of highly-customizable Venn and Euler diagrams in R. BMC Bioinformatics 2011; 12(1):35.
  • 23 Vieira GC. camposvieira/pln_cie_mrj: CIE [Internet]. Zenodo; 2025 [cited 2024 out 15]. Available from: https://doi.org/10.5281/zenodo.14747993
    » https://doi.org/10.5281/zenodo.14747993
  • 24 Morgan OW. How better pandemic and epidemic intelligence will prepare the world for future threats. Nat Med 2022; 28(8):1523-1526.
  • 25 Sahu KS, Majowicz SE, Dubin JA, Morita PP. NextGen Public Health Surveillance and the Internet of Things (IoT). Front Public Health 2021; 9:756675.
  • 26 Xu L, Zhou C, Luo S, Chan DK, McLaws ML, Liang W. Modernising infectious disease surveillance and an early-warning system: The need for China's action. Lancet Reg Health 2022; 23:100485.
  • 27 Seeskin ZH, LeClere F, Ahn J, Williams JA. Uses of Alternative Data Sources for Public Health Statistics and Policymaking: Challenges and Opportunities. JSM 2018 Govern Stat Section 2018; 1822-1861.
  • 28 Mit Critical Data. Secondary Analysis of Electronic Health Records [Internet]. Cham: Springer International Publishing; 2016 [cited 2024 set 30]. Available from: http://link.springer.com/10.1007/978-3-319-43742-2
    » http://link.springer.com/10.1007/978-3-319-43742-2
  • 29 Brasil. Ministério da Saúde (MS). Plano de ação para interrupção da circulação do vírus do sarampo: monitoramento e reverificação da sua eliminação no Brasil, 2022. Brasília: MS; 2022.
  • 30 Silva RPD, Pollettini JT, Pazin Filho A. Processamento de linguagem natural não supervisionado na identificação de pacientes suspeitos de infecção por COVID-19. Cad Saude Publica 2023; 39(11):e00243722.
  • 31 Ivanov O, Gesteland P, Hogan WR, Mundorff MB, Wagner M. Detection of Pediatric Respiratory and Gastrointestinal Outbreaks from Free-Text Chief Complaints. AMIA 2003 Symp Proc 2003; 318-322.
  • 32 Olszewski RT. Bayesian Classification of Triage Diagnoses for the Early Detection of Epidemics. Am Assoc Artif Intell 2003; 412-416.
  • 33 Organização Pan-Americana da Saúde (OPAS). As funções essenciais de saúde pública nas Américas - uma renovação para o século 21. Marco conceitual e descrição [Internet]. 2022 [acessado 2024 out 21]. Disponível em: https://iris.paho.org/handle/10665.2/55678
    » https://iris.paho.org/handle/10665.2/55678
  • Chief editors:
    Maria Cecília de Souza Minayo, Romeu Gomes, Antônio Augusto Moura da Silva

Publication Dates

  • Publication in this collection
    11 Aug 2025
  • Date of issue
    July 2025

History

  • Received
    24 Oct 2024
  • Accepted
    23 Jan 2025
  • Published
    25 Jan 2025
location_on
ABRASCO - Associação Brasileira de Saúde Coletiva Av. Brasil, 4036 - sala 700 Manguinhos, 21040-361 Rio de Janeiro RJ - Brazil, Tel.: +55 21 3882-9153 / 3882-9151 - Rio de Janeiro - RJ - Brazil
E-mail: cienciasaudecoletiva@fiocruz.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro