开放资源库 · 研究数据集

Vision-Language Models for Low-Vision Navigation Assistance in Cluttered Environments

Navigating busy, unstructured roads is a daily challenge for people with visual impairments, while many assistive systems do not explicitly account for the clutter and heterogeneous traffic found in Indian street environments. We present an end-to-end vision-to-audio navigation pipeline that transforms monocular video frames into priority-ordered spoken directional commands. The reported system combines a YOLOv8n detector fine-tuned on the Indian Driving Dataset (IDD), MiDaS DPT-LeViT monocular depth estimation, DeepSORT tracking, field-of-view-aware spatial reasoning, temporal smoothing, and an eight-layer rule-based navigation planner. The planner maps perceived hazards to concise commands such as STOP, AVOID, MOVE LEFT/RIGHT, and CONTINUE FORWARD, while a smart text-to-speech gate suppresses passive descriptions when an urgent command is pending. An optional BLIP-based visual validation stage is used in image mode. On the reported 60-frame manually annotated Indian-street evaluation set, the complete pipeline achieves 97.9% hazard recall, 92.8% distance accuracy, 0.407 walkable-corridor IoU, and 84.6 ms average latency per frame. On the IDD detection test set, fine-tuning improves mAP@50 from 0.389 for the COCO- pretrained baseline to 0.542, with mAP@50–95 increasing from 0.224 to 0.351. Ablations show that field-of-view-aware depth calibration is critical for proximity estimation, while temporal tracking is important for instruction stability and motion-aware warnings. The system is deliberately modular: the perception and spatial reasoning stack supplies grounded state information, while deterministic planning converts that state into actionable audio rather than relying on unconstrained generative narration.

← 返回资源筛选结果
RESOURCE OVERVIEW

资源说明

Navigating busy, unstructured roads is a daily challenge for people with visual impairments, while many assistive systems do not explicitly account for the clutter and heterogeneous traffic found in Indian street environments. We present an end-to-end vision-to-audio navigation pipeline that transforms monocular video frames into priority-ordered spoken directional commands. The reported system combines a YOLOv8n detector fine-tuned on the Indian Driving Dataset (IDD), MiDaS DPT-LeViT monocular depth estimation, DeepSORT tracking, field-of-view-aware spatial reasoning, temporal smoothing, and an eight-layer rule-based navigation planner. The planner maps perceived hazards to concise commands such as STOP, AVOID, MOVE LEFT/RIGHT, and CONTINUE FORWARD, while a smart text-to-speech gate suppresses passive descriptions when an urgent command is pending. An optional BLIP-based visual validation stage is used in image mode. On the reported 60-frame manually annotated Indian-street evaluation set, the complete pipeline achieves 97.9% hazard recall, 92.8% distance accuracy, 0.407 walkable-corridor IoU, and 84.6 ms average latency per frame. On the IDD detection test set, fine-tuning improves mAP@50 from 0.389 for the COCO- pretrained baseline to 0.542, with mAP@50–95 increasing from 0.224 to 0.351. Ablations show that field-of-view-aware depth calibration is critical for proximity estimation, while temporal tracking is important for instruction stability and motion-aware warnings. The system is deliberately modular: the perception and spatial reasoning stack supplies grounded state information, while deterministic planning converts that state into actionable audio rather than relying on unconstrained generative narration.

聚变工程工程验证Vision-Language ModelsLow-Vision AssistanceMonocular Depth EstimationVision-Based NavigationSpatial Reasoning
RESEARCH USE PROFILE

研究使用指引

适用任务诊断分析、模型校准、代理训练、跨装置比较与基准测试
使用准备核对字段、单位、缺失值、采样条件、训练测试划分和许可
核验重点需要检查数据泄漏、分布偏移以及装置和工况适用范围
RESOURCE FEEDBACK

资源信息需要更新?

可报告链接失效、文件异常、信息错误或版本变化,我们会核对并更新资源记录。