CHINESE JOURNAL OF ENERGETIC MATERIALS
+高级检索
面向含能材料知识图谱构建的大语言模型零样本信息抽取基准研究
作者:
作者单位:

1南京理工大学 化学与化工学院,江苏 南京 210094;2瞬态化学效应与控制全国重点实验室, 陕西 西安 710061;3南京理工大学 空间推进技术研究所, 江苏 南京 210094;4南京理工大学 微纳含能器件工业和信息化部重点实验室, 江苏 南京 210094;5中国科学院大连化学物理研究所人工智能中心, 辽宁 大连 116001

作者简介:

通讯作者:

基金项目:

瞬态化学效应与控制全国重点实验室稳定运行费资助项目 WDYX25614260207中央高校基本科研业务费专项资金 30926010601瞬态化学效应与控制全国重点实验室稳定运行费资助项目(WDYX25614260207);中央高校基本科研业务费专项资金(30926010601)


A Benchmark for Zero-shot Information Extraction of Large Language Models in Energetic Materials
Author:
Affiliation:

1Department of Applied Chemistry, School of Chemistry and Chemical Engineering, Nanjing University of Science and Technology, Nanjing 210094, China;2State Key Laboratory of Transient Chemical Effects and Control, Xi''an 710061, China;3Micro-nano Energetic Devices Key Laboratory of The Ministry of Industry and Information Technology, Nanjing 210094, China;4Institute of Space Propulsion, Nanjing University of Science and Technology, CAS, Nanjing 210094, China;5TER, The Dalian Institute of Chemical Physics, Dalian 116001, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 支撑附件
    摘要:

    从非结构化科学文献中准确提取材料实体、属性及实体-属性关联,是构建含能材料高质量数据集并为后续知识图谱构建提供结构化知识基础的关键环节。本研究构建中文含能材料全文信息抽取基准EM-Bench 1.0,选取2021~2026年国内核心期刊发表的40篇论文,标注材料名称、化学式、晶型、晶系、空间群、晶胞参数和密度,并将简化分子线性输入规范(SMILES)作为独立的探索性知识推理项。基准真值由两名领域研究人员独立标注,经交叉核验后对争议条目进行专家裁决。在任务层面零样本条件下,将每篇论文转换为完整Markdown文本,采用统一提示词和JSON输出格式,对ChatGPT、DeepSeek、Gemini和ChemELLM进行全文级评测。每个模型-文献组合在相同条件下独立运行3次,在有效输出中选取总体F1最高的一次作为最终评测结果,用于表征当前测试条件下大模型能够达到的较优表现。按材料-字段实例进行语义匹配和比较分析,Gemini总体F1最高(70.1%),ChatGPT召回率最高(82.2%);ChemELLM在材料名称(F1=70.5%)和晶型(F1=13.8%)抽取方面表现突出,晶胞参数单位缺失为0例,且抽取的无依据实体数最低(51个)。结果表明,现有大语言模型已具备一定的中文含能材料全文信息抽取能力,但仍存在漏抽、无依据生成和实体-属性错配等问题,尚不能在未经人工校核的情况下直接用于高可信知识入库。研究结果反映特定测试日期、网页配置和3次择优策略下的测试表现,不代表大模型的稳定性能或跨版本排名。EM-Bench 1.0以期为大语言模型选型与微调、含能材料高质量数据集构建以及面向知识图谱构建的自动化知识获取提供标准化评测依据。

    Abstract:

    Accurate extraction of material entities, attributes, and entity-attribute associations from unstructured scientific literature is a critical step in constructing high-quality datasets for energetic materials (EMs) and providing a structured knowledge foundation for subsequent knowledge graph construction. In this study, we developed EM-Bench 1.0, a full-text information extraction benchmark for Chinese-language literature on EMs. The benchmark comprises 40 articles published in major Chinese journals from 2021 to 2026 and includes annotations for material names, chemical formulas, polymorphs, crystal systems, space groups, unit-cell parameters, and densities. Simplified Molecular Input Line Entry System (SMILES) representations were treated separately as an exploratory knowledge-reasoning task. The benchmark ground truth was independently annotated by two domain researchers, followed by cross-validation, with disputed cases adjudicated by a domain expert. Under a task-level zero-shot setting, each article was converted into a complete Markdown document and evaluated using a unified prompt and JSON output format. Four large language model (LLM) platforms, ChatGPT, DeepSeek, Gemini, and ChemELLM, were evaluated at the full-text level. Each model-document pair was independently tested three times under identical conditions, and the valid output with the highest overall F1 score was retained as the final result to characterize the relatively strong performance attainable under the current test setting. Predictions were semantically matched and comparatively analyzed at the material-field instance level. Gemini achieved the highest overall F1 score (70.1%), whereas ChatGPT achieved the highest recall (82.2%). ChemELLM performed particularly well in material-name extraction (F1 = 70.5%) and polymorph extraction (F1 = 13.8%), produced no missing units for unit-cell parameters, and generated the fewest unsupported entities (51). The results indicate that current LLMs already possess a certain capability for full-text information extraction from Chinese literature on EMs; however, omissions, unsupported generations, and entity-attribute mismatches remain common. Therefore, model outputs should not be directly used for high-confidence knowledge ingestion without manual verification. The reported results reflect model performance under the specific test dates, web-interface configurations, and best-of-three selection strategy adopted in this study, rather than stable model performance or persistent cross-version rankings. EM-Bench 1.0 is intended to provide a standardized benchmark for large language model selection and fine-tuning, EM high-quality dataset construction, and automated knowledge acquisition for knowledge graph construction.

    参考文献
    相似文献
    引证文献
文章指标
  • PDF下载次数:
  • HTML阅读次数:
  • 摘要点击次数:
  • 引用次数:
引用本文

臧小为,钱志翔,葛一帆,等. 面向含能材料知识图谱构建的大语言模型零样本信息抽取基准研究[J]. 含能材料,DOI:10.11943/CJEM2026166.

复制
历史
  • 收稿日期: 2026-07-13
  • 最后修改日期: 2026-10-04
  • 录用日期: 2026-09-08
  • 在线发布日期: 2026-10-09
  • 出版日期: