{"database":"biostudies-literature","file_versions":[],"scores":null,"additional":{"omics_type":["Unknown"],"volume":["9(1)"],"submitter":["Anjos de Almeida V"],"pubmed_abstract":["<h4>Objectives</h4>Medical coding structures health-care data for research, quality monitoring, and policy. This study assesses the potential of large language models (LLMs) to assign International Classification of Primary Care, 2nd edition (ICPC-2) codes using the output of a domain-specific search engine.<h4>Materials and methods</h4>A dataset of 437 Brazilian Portuguese clinical expressions, each annotated with ICPC-2 codes, was used. A semantic search engine (OpenAI's text-embedding-3-large) retrieved candidates from 73 563 labeled concepts. Thirty-three LLMs were prompted with each query and retrieved results to select the best-matching ICPC-2 code. Performance was evaluated using F1-score, along with token usage, cost, response time, and format adherence.<h4>Results</h4>Twenty-eight"],"journal":["JAMIA open"],"pagination":["ooag017"],"full_dataset_link":["https://www.ebi.ac.uk/biostudies/studies/S-EPMC12924630"],"repository":["biostudies-literature"],"pubmed_title":["Large language models as medical code selectors: a benchmark using the International Classification of Primary Care."],"pmcid":["PMC12924630"],"pubmed_authors":["de Camargo V","Gomez-Bravo R","Fernandez Lopez L","Finger M","van der Haring E","Anjos de Almeida V","van Boven K"],"additional_accession":[]},"is_claimable":false,"name":"Large language models as medical code selectors: a benchmark using the International Classification of Primary Care.","description":"<h4>Objectives</h4>Medical coding structures health-care data for research, quality monitoring, and policy. This study assesses the potential of large language models (LLMs) to assign International Classification of Primary Care, 2nd edition (ICPC-2) codes using the output of a domain-specific search engine.<h4>Materials and methods</h4>A dataset of 437 Brazilian Portuguese clinical expressions, each annotated with ICPC-2 codes, was used. A semantic search engine (OpenAI's text-embedding-3-large) retrieved candidates from 73 563 labeled concepts. Thirty-three LLMs were prompted with each query and retrieved results to select the best-matching ICPC-2 code. Performance was evaluated using F1-score, along with token usage, cost, response time, and format adherence.<h4>Results</h4>Twenty-eight","dates":{"release":"2026-01-01T00:00:00Z","publication":"2026 Feb","modification":"2026-07-09T12:08:56.656Z","creation":"2026-07-09T11:10:12.375Z"},"accession":"S-EPMC12924630","cross_references":{"pubmed":["41727414"],"doi":["10.1093/jamiaopen/ooag017"]}}