Skip to content
#

bertscore

Here are 47 public repositories matching this topic...

LLM evaluation on a 643-question Austrian tax-law benchmark: ROUGE, BLEU and BERTScore plus a manual failure-mode analysis. Compares a QLoRA fine-tuned 5B model against a 31B zero-shot baseline; the smaller fine-tuned model wins on all five metrics.

  • Updated Jul 26, 2026
  • Jupyter Notebook

Implementation of a task-specific QLoRA supervised fine-tuning pipeline for LLaMA-2-7B-Chat, developed for an independent study on structured cover letter generation.

  • Updated Jun 9, 2026
  • Python

The work presented was developed during the internship, as researchers in the field of Natural Language Generation, at the Insid&s Lab laboratory in Milan-Bicocca. The work carried out deals with the creation of a framework for the correct assessment of the impact of the quality of the input datasets on the quality of the text generated by the N…

  • Updated Mar 30, 2026
  • Jupyter Notebook
radscore

[radscore] Evaluation metrics for radiology report generation — BLEU, ROUGE, BERTScore, F1-RadGraph, CheXbert F1 & GREEN in one `pip install`.

  • Updated Jun 16, 2026
  • Python

An end-to-end prompt-based navigation task for the Pioneer Valley region: synthetic data, a custom evaluation suite, and several methods helping map the problem's Pareto front

  • Updated Jul 4, 2026
  • Jupyter Notebook

Improve this page

Add a description, image, and links to the bertscore topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the bertscore topic, visit your repo's landing page and select "manage topics."

Learn more