Biomedical LLM Hallucination Detection via Classifier-Conditioned Factuality Classification
Students & Supervisors
Student Authors
Supervisors
Abstract
General-purpose LLMs have been shown to have a strong capability to generate coherent texts in the biomedical domain with a high level of confidence. How- ever, the output texts are often factually inaccurate or unsupported. Such hallucinations pose risks in high-stakes applications. This paper proposes a supervised framework for hallucination detection using a fine-tuned PubMedBERT classifier for binary factuality classification. The experimental results show that the proposed framework performs exceptionally well in discriminating factuality. The accuracy achieved is 98%, and the precision, recall, and F1-score achieved are 97.22%, 99.05%, and 98.13%, respectively. The model also demonstrates a stable ROC-AUC. The framework’s stability is further demonstrated by a threshold sensitivity analysis. Threshold sensitivity analysis (0.3–0.7) shows no sudden performance degradation, indicating robustness. The approach serves as an effective verification mechanism for LLM-generated biomedical content. Its reliance on expert-validated supervision, combined with lightweight retrieval grounding applied only during inference, enables a latency-efficient verification mechanism
Keywords
Publication Details
- Type of Publication:
- Conference Name: International Conference on Electrical, Computer and Communication Technologies (ECCT 2026)
- Date of Conference: 05/07/2026 - 05/07/2026
- Venue: Dhaka International University, Bangladesh
- Organizer: Prof. Dr. Md. Abdul Based