CHINESE JOURNAL OF ENERGETIC MATERIALS
+Advanced Search

A Benchmark for Zero-shot Information Extraction of Large Language Models in Energetic Materials
Author:
Affiliation:

1Department of Applied Chemistry, School of Chemistry and Chemical Engineering, Nanjing University of Science and Technology, Nanjing 210094, China;2State Key Laboratory of Transient Chemical Effects and Control, Xi''an 710061, China;3Micro-nano Energetic Devices Key Laboratory of The Ministry of Industry and Information Technology, Nanjing 210094, China;4Institute of Space Propulsion, Nanjing University of Science and Technology, CAS, Nanjing 210094, China;5TER, The Dalian Institute of Chemical Physics, Dalian 116001, China

Fund Project:

  • Article
  • |
  • Figures
  • |
  • Metrics
  • |
  • Reference
  • |
  • Related
  • |
  • Cited by
  • |
  • Materials
    Abstract:

    Accurate extraction of material entities, attributes, and entity-attribute associations from unstructured scientific literature is a critical step in constructing high-quality datasets for energetic materials (EMs) and providing a structured knowledge foundation for subsequent knowledge graph construction. In this study, we developed EM-Bench 1.0, a full-text information extraction benchmark for Chinese-language literature on EMs. The benchmark comprises 40 articles published in major Chinese journals from 2021 to 2026 and includes annotations for material names, chemical formulas, polymorphs, crystal systems, space groups, unit-cell parameters, and densities. Simplified Molecular Input Line Entry System (SMILES) representations were treated separately as an exploratory knowledge-reasoning task. The benchmark ground truth was independently annotated by two domain researchers, followed by cross-validation, with disputed cases adjudicated by a domain expert. Under a task-level zero-shot setting, each article was converted into a complete Markdown document and evaluated using a unified prompt and JSON output format. Four large language model (LLM) platforms, ChatGPT, DeepSeek, Gemini, and ChemELLM, were evaluated at the full-text level. Each model-document pair was independently tested three times under identical conditions, and the valid output with the highest overall F1 score was retained as the final result to characterize the relatively strong performance attainable under the current test setting. Predictions were semantically matched and comparatively analyzed at the material-field instance level. Gemini achieved the highest overall F1 score (70.1%), whereas ChatGPT achieved the highest recall (82.2%). ChemELLM performed particularly well in material-name extraction (F1 = 70.5%) and polymorph extraction (F1 = 13.8%), produced no missing units for unit-cell parameters, and generated the fewest unsupported entities (51). The results indicate that current LLMs already possess a certain capability for full-text information extraction from Chinese literature on EMs; however, omissions, unsupported generations, and entity-attribute mismatches remain common. Therefore, model outputs should not be directly used for high-confidence knowledge ingestion without manual verification. The reported results reflect model performance under the specific test dates, web-interface configurations, and best-of-three selection strategy adopted in this study, rather than stable model performance or persistent cross-version rankings. EM-Bench 1.0 is intended to provide a standardized benchmark for large language model selection and fine-tuning, EM high-quality dataset construction, and automated knowledge acquisition for knowledge graph construction.

    Reference
    Related
    Cited by
Article Metrics
  • PDF:
  • HTML:
  • Abstract:
  • Cited by:
Get Citation

ZANG xiaowei, QIAN zhixiang, GE yifan, et al. A Benchmark for Zero-shot Information Extraction of Large Language Models in Energetic Materials[J]. Chinese Journal of Energetic Materials(Hanneng Cailiao),DOI:10.11943/CJEM2026166.

Cope
History
  • Received:July 13,2026
  • Revised:October 04,2026
  • Adopted:September 08,2026
  • Online: October 09,2026
  • Published: