摘要

背景:准确术前评估下颌第三磨牙(M3Ms)与下牙槽神经管(IAC)的空间关系对于最大限度地降低神经损伤风险至关重要。虽然全景片(PRs)通常用作一线模态,但其二维性质限制了诊断可靠性,通常需要确认性锥形束CT(CBCT)。

方法:我们开发了一种双流、Transformer增强的多模态(全景片 + 选定的CBCT切片)深度学习框架,设计为“临床医生在环”系统。该模型联合处理配对的全景片和操作员选择的CBCT切片,以模拟现实的专家咨询。在250例配对PR-CBCT病例的数据集上系统地评估了四种广泛使用的高性能骨干网络(ConvNeXt、Swin Transformer、ResNet-50和VGG16)。使用F1分数、精确度、召回率和ROC曲线下面积(AUC)评估模型性能。使用基于SHAP的可解释性分析探索模型可解释性。

结果:在评估的架构中,基于ConvNeXt的模型实现了最高的整体性能,宏F1分数为0.9478,精确度为0.9515,召回率为0.9468。学习的嵌入显示出良好的类别可分性,可解释性分析突出了与既定放射学风险征象一致的具有解剖学意义的区域。

结论:提出的多模态2D框架展示了具有竞争力且稳健的分类性能。通过将专家切片选择与自动特征提取相结合,该模型作为一个标准化的决策支持工具发挥作用,旨在减少在单独视觉解释可能主观的复杂边界病例中的诊断变异性。

原文摘要

BACKGROUND: Accurate preoperative assessment of the spatial relationship between mandibular third molars (M3Ms) and the inferior alveolar canal (IAC) is essential to minimize the risk of nerve injury. While panoramic radiographs (PRs) are routinely used as a first-line modality, their two-dimensional nature limits diagnostic reliability, often necessitating confirmatory cone-beam computed tomography (CBCT).

METHODS: We developed a dual-stream, transformer-enhanced multimodal (panoramic radiography + selected CBCT slices) deep learning framework designed as a clinician-in-the-loop system. The model jointly processes paired panoramic radiographs and operator-selected CBCT slices to simulate a realistic specialist consultation. Four widely used high-performance backbones (ConvNeXt, Swin Transformer, ResNet-50, and VGG16) were systematically evaluated on a dataset of 250 paired PR-CBCT cases. Model performance was assessed using F1-score, precision, recall, and area under the ROC curve (AUC). Model interpretability was explored using SHAP-based explainability analysis.

RESULTS: Among the evaluated architectures, the ConvNeXt-based model achieved the highest overall performance, with a macro F1-score of 0.9478, precision of 0.9515, and recall of 0.9468. The learned embeddings showed good class separability, and explainability analyses highlighted anatomically meaningful regions consistent with established radiographic risk signs.

CONCLUSION: The proposed multimodal 2D framework demonstrates competitive and robust classification performance. By integrating expert slice selection with automated feature extraction, the model functions as a standardized decision-support tool, aimed at reducing diagnostic variability in complex boundary cases where visual interpretation alone may be subjective.

出处

BMC oral health 2026;26(1). DOI: 10.1186/s12903-026-09012-z.

PubMed 原文链接(PMID 42668366)

本页仅收录摘要与出处,不存储全文(合规要求)。