基于多模态视觉自注意力机制的机器人目标识别研究
DOI:
CSTR:
作者:
作者单位:

郑州大学

作者简介:

通讯作者:

中图分类号:

基金项目:

郑州市社会科学调研课题 项目编号:ZSJX20250794


Research on Robot Target Recognition Based on Multimodal Visual Self-Attention Mechanism
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对现有机器人目标识别算法存在的目标识别精确率低、虚警率过高等不足,在多模态条件下设计一种视觉自注意力机制网络模型。先基于动态窗口设计了机器人的运动模型,采集多模态数据并数据进行配准预处理,基于评价函数评估机器人运动姿态精度。然后引入了CNN网络模型作为注意力模块的骨干网络,提升注意力机制网络模型的训练能力和数据处理能力,通过位置编码提取关键特征信息,最后基于双门控机制进一步提升机器人目标识别算法的识别精度和特征提取能力。实验测试结果显示,在无遮挡和有遮挡两种情况下提出识别算法测试集的识别精确率、虚警率分别为99.1%,0.4%,98.0%,1.0%,均显著优于现有的识别算法。通过消融实验验证了模型的多次改进对模型性能提升的贡献率。

    Abstract:

    Aiming at the shortcomings of the existing robot target recognition algorithms, such as low target recognition accuracy and excessive false alarm rate, a visual self-attention mechanism network model is designed under multimodal conditions. Firstly, the motion model of the robot was designed based on the dynamic window. Multimodal data was collected and the data was registered and preprocessed. The motion posture accuracy of the robot was evaluated based on the evaluation function. Then, the CNN network model was introduced as the backbone network of the attention module to enhance the training ability and data processing ability of the attention mechanism network model. Key feature information was extracted through position encoding. Finally, based on the gating mechanism, the recognition accuracy and feature extraction ability of the robot target recognition algorithm were further improved. The experimental test results show that the recognition accuracy rate and false alarm rate of the proposed recognition algorithm test set in both unoccluded and occluded situations are 99.1%, 0.4%, 98.0%, and 1.0% respectively, which are significantly better than the existing recognition algorithms. The contribution rate of multiple improvements to the model"s performance enhancement was verified through ablation experiments.

    参考文献
    相似文献
    引证文献
引用本文

马文静,李亚哲,高军亮.基于多模态视觉自注意力机制的机器人目标识别研究计算机测量与控制[J].,2026,34(7):261-267.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-12-04
  • 最后修改日期:2026-03-20
  • 录用日期:2026-01-16
  • 在线发布日期: 2026-07-24
  • 出版日期:
文章二维码