Helmholtz Gemeinschaft

Search
Browse
Statistics
Feeds

HunFlair: an easy-to-use tool for state-of-the-art biomedical named entity recognition

[thumbnail of Original Article]
Preview
PDF (Original Article) - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
161kB
[thumbnail of Supplementary Material]
Preview
PDF (Supplementary Material) - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
223kB

Item Type:Article
Title:HunFlair: an easy-to-use tool for state-of-the-art biomedical named entity recognition
Creators Name:Weber, L., Sänger, M., Münchmeyer, J., Habibi, M., Leser, U. and Akbik, A.
Abstract:SUMMARY: Named entity recognition (NER) is an important step in biomedical information extraction pipelines. Tools for NER should be easy to use, cover multiple entity types, be highly accurate, and be robust towards variations in text genre and style. We present HunFlair, an NER tagger fulfilling these requirements. HunFlair is integrated into the widely-used NLP framework Flair, recognizes five biomedical entity types, reaches or overcomes state-of-the-art performance on a wide set of evaluation corpora, and is trained in a cross-corpus setting to avoid corpus-specific bias. Technically, it uses a character-level language model pretrained on roughly 24 million biomedical abstracts and three million full texts. It outperforms other off-the-shelf biomedical NER tools with an average gain of 7.26 pp over the next best tool in a cross-corpus setting and achieves on-par results with state-of-the-art research prototypes in in-corpus experiments. HunFlair can be installed with a single command and is applied with only four lines of code. Furthermore, it is accompanied by harmonized versions of 23 biomedical NER corpora. AVAILABILITY: HunFlair ist freely available through the Flair NLP framework (https://github.com/flairNLP/flair) under an MIT license and is compatible with all major operating systems. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Source:Bioinformatics
ISSN:1367-4803
Publisher:Oxford University Press
Volume:37
Number:17
Page Range:2792–2794
Date:1 September 2021
Additional Information:Copyright © The Author(s) 2021. Published by Oxford University Press.
Official Publication:https://doi.org/10.1093/bioinformatics/btab042
External Fulltext:View full text on PubMed Central
PubMed:View item in PubMed

Repository Staff Only: item control page

Downloads

Downloads per month over past year

Open Access
MDC Library