How Transparência Brasil Uses Artificial Intelligence to Identify Medications in Public Procurement

The platform cross-references PNCP descriptions with the federal catalog and achieves 98% accuracy in item classification, enabling price monitoring in the healthcare sector
Data de publicação
02/06/2026
Public procurement

The “Transparent Medication Price Basket,” a price search platform developed by Transparência Brasil to assist managers and civil servants in the efficient procurement of public medications, used artificial intelligence to solve a problem that had made any oversight impossible: although the National Public Procurement Portal (PNCP) published data on approximately 12 million items purchased in 2024 alone, there was no mechanism to distinguish medications from other procured goods.

The PNCP centralizes procurement by public entities across the country and, since the passage of the New Law on Bidding and Contracts (14.133/2021), has become the primary instrument for transparency in this area. The same law requires purchasers to consult the Health Price Database—a database containing reference prices for medications and health products—when making purchases for the universal healthcare system. However, without identifying which items are medications, this requirement cannot be monitored. And there is no way to make this identification manually in a database of this size: it would take a human about 90 minutes to analyze 100 items by consulting a catalog.

The problem is exacerbated by the lack of standardization in the portal’s descriptions. Items are recorded in free-form text, without following a mandatory standard or catalog. Medications appear with abbreviations, inconsistent spellings, and missing accents, or with the product name, concentration, and packaging information crammed into a single field. A description such as “Sodium Dipyrone, Dosage: 500 mg – Tablet,” for example, may be entered into the PNCP as “Dipyrone S. 500 mg tab.”

How It Works

The first step was to match the item descriptions in the public procurement system—using open-contract data published in the PNCP—with the entries in the federal catalog of goods and services, CATMAT. 

In the PNCP, descriptions of procured goods and services appear as free-form text in the item fields associated with bids and contracts. We then used an LLM-based embedding model to associate each PNCP item related to medications with a CATMAT entry (identified by its catalog code, “Br code”), based on the similarity between the descriptions. Next, a similarity threshold was applied to determine whether the item should be classified as a medication.

Results

When tested on a sample of 1,000 manually labeled items, the model achieved 98% accuracy in classifying items as either medications or non-medications, and 86% accuracy in identifying the correct Br code in the national catalog. Accuracy can be assessed by calculating the precision, recall, and overall accuracy of the model’s predictions.

Reproducible

The project was developed as part of the Lift impact accelerator program of the Open Contracting Partnership (OCP), in partnership with the Secretariat of Management and Innovation of the Ministry of Management and Innovation and the Office of the Comptroller General. The source code is publicly available on Transparência Brasil’s GitHub. 

The approach is replicable: any procurement system that has free-text item descriptions and a reference catalog can benefit from a similar solution, whether to monitor medications, compare prices, or expand oversight in other sectors.

Apoie a transparência dos dados públicos