Skip navigation
Abstract Cube Background With Red Lights And Programming Code
SEI brief

Classifying and benchmarking LLM-extracted summaries for policy evaluation

Start reading
SEI brief

Classifying and benchmarking LLM-extracted summaries for policy evaluation

This is the second brief in a series of three on how SEI is testing methods for AI-enabled systematic literature review and evidence synthesis.

Adis Dzebo, Kira Kappe / Published on 18 August 2026

Citation

Dzebo, A., Kappe, K. (2026). Classifying and benchmarking LLM-extracted summaries for policy evaluation. (SEI Brief). https://doi.org/10.51414/sei2026.033

1. Introduction

This is the second of three briefs on methods for AI-enabled systematic literature review and evidence synthesis. In the first, we discuss how large language models (LLMs), as well as SEI’s AI Reader, could support information extraction from unstructured documents at scale and the steps that need to be taken to ensure reliable and accurate results (Kappe & Dzebo, 2025).

The SEI AI Reader demonstrated that LLMs can extract analytically usable summaries from unstructured evaluation documents at approximately 85% agreement with human coding. However, extraction alone does not yield comparable data.

This, the second brief in the series, explains how the resulting 111 000 summaries, extracted from 2839 policy evaluations and independent audits, are converted into a comparable evidence base for detailed analysis of the processes and conditions that were responsible for policy success and failure in 124 countries and across 21 sectors. This data and its classification are the major component of the Outcomes and Process Evidence from Reviews and Audits (OPERA) database (Dzebo et al., 2026).

2. Devising a structure for analytical processing

The SEI AI Reader (Babis et al., 2024) is designed to enable a structured analysis where dependent and independent variables are defined, and additional context provides flexibility to iteratively fine-tune through variable-specific input and guidelines.

The policy evaluation framework applied in the first brief was continuously used throughout the AI-processing pipeline. It is structured around three layers. Benchmark variables (Part 0) capture policy objectives, success criteria, theory of change, evaluation methods, and main findings, anchoring what each evaluation set out to assess and how rigorously it did so. Outcome variables (Part 1) cover effectiveness, efficiency, outcomes and impact, and attribution, capturing what was achieved. Process variables (Part 2) cover agenda-setting, formulation, content coherence, implementation, and engagement and inclusion, capturing how it was achieved. This separation enables us to study process-outcome relationships in policy evaluation.

A pre-processing step provided a decisive improvement in data quality. It included text cleaning (e.g. removing headers and footers and margin text) and processing all documents from PDF to AI-friendly Markdown format using a tailored optical character recognition (OCR) process that extracted tables and figures separately from the main document, using three separate models,1 and merging them together in a post-processing step.

This step increased mean summaries per document from 21 to 35 and provided a balanced distribution between outcome and process summaries. The extracted data is summarized in Table 1.

Table 1. Database overview

Metric Database value
Total documents 2839
Total summaries 111 436
Mean summaries per document 39.3
Countries 124
Sectors 21
Benchmark summaries 13 759
Outcome summaries 52 841
Process summaries 44 836

3. Variable classification with natural language processing

Using LLMs for repeated classification of text would incur significant financial costs as well as contribute to unnecessary environmental impact due to their energy consumption. Results would also be non-deterministic because the same input would produce different outputs at different runs (Atil et al., 2025). Instead, a lighter sentence transformer an embedding model was used to encode as a 768-dimensional vector and passed to a logistic regression classifier that was trained for each of the nine variables (Bui et al., 2016).

The BAAI/bge-base-en-v1.5 base embedding model was fine-tuned on 14 739 contrastive text pairs drawn from manually labelled summaries, using “multiple negatives ranking loss” (Hoffmann et al., 2022). This process of domain fine-tuning adapts the model’s vector space to evaluative language, improving its ability to distinguish phrases such as “achieved 40% of the target” from “achieved 90% of the target”, which generic models treat as near-identical (Zhang et al., 2025).

The OPERA database at a glance: five steps from text to evaluation score

  1. The SEI AI Reader extracts 111 436 analytical summaries from 2839 documents.
  2. Nine fine-tuned natural language processing (NLP) classifiers assign all summaries a classification based on its variable belonging and manual validation, achieving production-grade accuracy.
  3. Five cross-variable gates remove inconsistent and/or weak confidence predictions.
  4. A combination of basis score and bonus score for outcome and process variable summaries aggregates to a document score
  5. Summaries are aggregated to a 0–100 score and assigned to one of five success tiers.

4. Model training and inference

Systematic A/B testing of both researcher-validated summaries and random database examples compared the base model against the fine-tuned. This resulted in a hybrid combination where six fine-tuned variable models performed better than the base model. Three variables – attribution, outcomes and impact, and content coherence – which depend on auxiliary information rather than evaluative signals, performed worse after fine-tuning. These variables depend on contextual and definitional reasoning, for example recognizing what constitutes a strong causal claim, or whether a policy aligns with international framework mechanisms.  

A parallel ablation process showed that the justification text, an LLM-generated motivation summary for each extraction was shown to improve accuracy between 6 and 11 percentage points. In contrast, similar tests with benchmark summaries gave mixed results, mainly due to crowding of the 512-token context window of the model.

This combination produced mean production test accuracy of 90% across the nine variables, ranging from 81.1% (Effectiveness) to 100% (Agenda-setting and content coherence). Table 2 summarizes the performance for all nine variables.

Table 2. Classification of model performance

Variable Model* Test accuracy CV accuracy Precision Recall F1 Total n
1. Effectiveness FT 81.1% 73.0% 0.81 0.81 0.81 481
2. Efficiency FT 80.0% 96.0% 0.81 0.81 0.84 477
3. Outcomes and impact Base 97.0% 89.2% 0.97 0.97 0.97 262
4. Attribution Base 89.1% 77.4% 0.91 0.89 0.89 507
5. Agenda-setting FT 100% 96.2% 1.00 1.00 1.00 291
6. Formulation FT 90.2% 81.2% 0.92 0.90 0.90 445
7. Content coherence Base 100% 98.0% 1.00 1.00 1.00 605
8. Implementation FT 92.0% 82.3% 0.95 0.95 0.95 330
9. Engagement FT 80.8% 77.4% 0.82 0.81 0.80 314

*FT= fine-tuned
Note 1: Test accuracy on held-out 20% stratified test set.
Note 2: Models were trained in one-stage, two-stage or two-phase processes, which were determined by the number of extracted summaries that were either multi-tagged across variables or classified as “no data”. Table 2 shows the values of the last step for each variable.

5. Refining and quality control

Two additional post-training and post inference steps were applied for increased reliability. First, each prediction carries a probability that is assigned one of four confidence levels – Strong, Moderate, Weak, or Unclear – and is based on the manually classified summaries. Thresholds were adjusted and calibrated per variable, matching the model accuracy. For example, a variable with high accuracy was allowed a relaxed threshold for its top label, whereas a less accurate one requires higher probability before its prediction is trusted. In addition, four variables applied temperature scaling post hoc to correct miscalibrated probabilities (Guo et al., 2017). Calibration of confidence levels, in turn, weighted each summary’s total contribution to the aggregated score, ensuring that low-quality predictions are attenuated rather than discarded.

Table 3. Variable categories with summary scores

Variable Top-tier score Mid-tier score Bottom score Negative score
1. Effectiveness Effective (90) Partially Effective (60) Not effective (10)
2. Efficiency Efficient (90)   Not efficient (20) Inefficient (5)
3. Outcome and impact Outcome and impact (95) No outcome (20) Negative outcome or impact (5)
4. Attribution Strong attribution (95) Weak or no attribution (20) Negative attribution (5)
5. Agenda setting Strong agenda-setting (90) Mixed (60) Conflictive (20)
6. Formulation Well-drafted (90) Adequately Drafted (60) Poorly drafted (10)
7. Content coherence Coherent (90) Not coherent (20) Incoherent (5)
8. Implementation Effective (90) Ineffective (5)
9. Engagement Meaningful (90) Not meaningful (20) Insufficient (5)

Second, cross-variable consistency gates that applied the original predictions and adjusted them (reclassified or downgraded) based on document-level evidence patterns. For example, our reasoning was that a policy must have a positive outcome before its causality can be assessed. Therefore, a document classified as “Not Effective” (i.e. where the Variable 1 Effectiveness summaries were overwhelmingly classified as “Not Effective”, or no summaries were classified as “Effective”) cannot credibly carry “Strong Attribution”. In such cases, “Strong Attribution” was downgraded to “Weak or No Attribution”.

In addition, four other gates were applied:

  • “Outcome and impact” was changed to “negative outcome and impact” when the benchmark classification “main findings” and the variable “Effectiveness” both independently indicated failure.
  • Variable 8 gate cross-validates “effective implementation” (Variable 8) against “outcome evidence” (Variable 3).
  • Efficiency gate cross-validates efficient predictions against “effectiveness” (variable 1) and “outcomes” (variable 3).
  • Confidence-downgrade gate, which demotes all summaries with weak or unclear confidence.

The categories of variables and summary scores are presented in Table 3.

6. Aggregating summaries for document-level evaluation

A simple average of summary classifications is inadequate for producing comparable document scores, because of the varying number of summaries per evaluation and the asymmetry between four outcome variables and five process variables. Instead, we developed a scoring methodology following established composite indicator methodology (Greco et al., 2019; OECD et al., 2008).

The methodology consists of four scoring components. The first two include individual basis scores for outcome and process variables. For outcomes, each of the four variables contributes up to 12.5 points, decided by the highest scoring variable summary. For process, with five variables, each variable could contribute 10 points. Here, a breadth-first approach was applied to ensure that all variables were included when available. In both cases, the maximum basis score was 50 points.

Second, two quality-bonus methods complemented the outcome and process basis scores. A bonus was awarded only when at least two top-tier summaries co-occur across multiple variables. Bonuses increased with the number of top-tier summaries per variable and could add 10 points in total for both outcome and process basis scores.

A dampening effect was applied based on the mean confidence, as well as for long documents with a larger amount of summaries, using effective sample size methodology (Kish, 1965). Finally, despite the theoretical maximum score being 120, scores were capped at 100, and a mid-tier ceiling was put on those evaluations that did not have any top-tier classified summaries.

The resulting score (0–100) was categorized in five success tiers: Elite (≥ 90), Gold (75–89), Good (50–74), Moderate (25–49), and Poor (< 25). The final tier distribution of the 2839 evaluations documents and other major findings from the OPERA database are documented in the third brief in the series.

7. Conclusion

This brief introduced the text classification methodology, combining policy evaluation frameworks, policy implementation theory and natural language processing (NLP), used in the OPERA database (Dzebo et al., 2026). The methodology was applied to more than 111 000 summaries extracted from 2839 policy evaluations and independent audits. The scoring methodology is deliberately optimized for identifying well-evidenced success. Below the top two tiers, low scores may reflect genuinely unsuccessful policies, incomplete reporting, or both. To address the complementary question of why policies fail, a separate failure-mechanism classification was developed and will be discussed in forthcoming publications.

References

Atil, B., Aykent, S., Chittams, A., Fu, L., Passonneau, R. J., Radcliffe, E., Rajagopal, G. R., Sloan, A., Tudrej, T., Ture, F., Wu, Z., Xu, L., & Baldwin, B. (2025). Non-Determinism of “Deterministic” LLM Settings (arXiv:2408.04667). arXiv. https://doi.org/10.48550/arXiv.2408.04667

Babis, W., Munoz Cabre, M., Dzebo, A., Martelo, C., Salzano, C., Torres Morales, E., & Arsadita, F. (2024). SEI AI Policy Reader [Dataset]. Stockholm Environment Institute. https://gptbatchpolicyprocesstool.streamlit.app/

Bui, D. D. A., Del Fiol, G., & Jonnalagadda, S. (2016). PDF text classification to leverage information extraction from publication reports. Journal of Biomedical Informatics, 61, 141–148. https://doi.org/10.1016/j.jbi.2016.03.026

Clark, C., & Divvala, S. (2016). PDFFigures 2.0: Mining Figures from Research Papers. JCDL. http://pdffigures2.allenai.org/

Dzebo, A., Kappe, K., Babis, W., & Cabre, M. M. (2026). Outcomes and Process Evidence from Reviews and Audits (OPERA) Database (Version 1.0) [Dataset]. https://polevaldata.kolle44server.org/

Greco, S., Ishizaka, A., Tasiou, M., & Torrisi, G. (2019). On the Methodological Framework of Composite Indices: A Review of the Issues of Weighting, Aggregation, and Robustness. Social Indicators Research, 141(1), 61–94. https://doi.org/10.1007/s11205-017-1832-9

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning – Volume 70, ICML’17, 1321–1330. https://dl.acm.org/doi/10.5555/3305381.3305518

Hoffmann, D. T., Behrmann, N., Gall, J., Brox, T., & Noroozi, M. (2022). Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked Positives. Proceedings of the AAAI Conference on Artificial Intelligence, 36(1), 897–905. https://doi.org/10.1609/aaai.v36i1.19972

Kappe, K., & Dzebo, A. (2025). Can AI produce reliable and consistent data analysis? https://doi.org/10.51414/sei2025.040

Kish, L. (1965). Sampling Organizations and Groups of Unequal Sizes. American Sociological Review, 30(4), 564–572. https://doi.org/10.2307/2091346

Li, Z., Liu, Y., Liu, Q., Ma, Z., Zhang, Z., Zhang, S., Guo, Z., Zhang, J., Wang, X., & Bai, X. (2025). MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm (arXiv:2506.05218). arXiv. https://doi.org/10.48550/arXiv.2506.05218

OECD, European Union, & European Commission – Joint Research Centre. (2008). Handbook on Constructing Composite Indicators: Methodology and User Guide. OECD. https://doi.org/10.1787/9789264043466-en

Smock, B., Pesala, R., & Abraham, R. (2021). PubTables-1M: Towards comprehensive table extraction from unstructured documents (arXiv:2110.00061). arXiv. https://doi.org/10.48550/arXiv.2110.00061

Zhang, Y., Zhu, Y., Zhu, Z., Liu, P., Xie, P., & Wu, C. (2025). A Domain-Finetuned Semantic Matching Framework Based on Dynamic Masking and Contrastive Learning for Specialized Text Retrieval. Electronics, 14(24). https://doi.org/10.3390/electronics14244882

Notes

SEI authors

Adis Dzebo
Adis Dzebo

Senior Research Fellow

SEI Headquarters

Kira Kappe
Kira Kappe

Research Associate

SEI Headquarters