Vision-Language Models for Low-Vision Navigation Assistance in Cluttered Environments
Navigating busy, unstructured roads is a daily challenge for people with visual impairments, while many assistive systems do not explicitly account for the clutter and heterogeneous traffic found in Indian street environments. We present an end-to-end vision-to-audio navigation pipeline that transforms monocular video frames into priority-ordered spoken directional commands. The reported system combines a YOLOv8n detector fine-tuned on the Indian Driving Dataset (IDD), MiDaS DPT-LeViT monocular depth estimation, DeepSORT tracking, field-of-view-aware spatial reasoning, temporal smoothing, and an eight-layer rule-based navigation planner. The planner maps perceived hazards to concise commands such as STOP, AVOID, MOVE LEFT/RIGHT, and CONTINUE FORWARD, while a smart text-to-speech gate suppresses passive descriptions when an urgent command is pending. An optional BLIP-based visual validation stage is used in image mode. On the reported 60-frame manually annotated Indian-street evaluation set, the complete pipeline achieves 97.9% hazard recall, 92.8% distance accuracy, 0.407 walkable-corridor IoU, and 84.6 ms average latency per frame. On the IDD detection test set, fine-tuning improves mAP@50 from 0.389 for the COCO- pretrained baseline to 0.542, with mAP@50–95 increasing from 0.224 to 0.351. Ablations show that field-of-view-aware depth calibration is critical for proximity estimation, while temporal tracking is important for instruction stability and motion-aware warnings. The system is deliberately modular: the perception and spatial reasoning stack supplies grounded state information, while deterministic planning converts that state into actionable audio rather than relying on unconstrained generative narration.
资源说明
Navigating busy, unstructured roads is a daily challenge for people with visual impairments, while many assistive systems do not explicitly account for the clutter and heterogeneous traffic found in Indian street environments. We present an end-to-end vision-to-audio navigation pipeline that transforms monocular video frames into priority-ordered spoken directional commands. The reported system combines a YOLOv8n detector fine-tuned on the Indian Driving Dataset (IDD), MiDaS DPT-LeViT monocular depth estimation, DeepSORT tracking, field-of-view-aware spatial reasoning, temporal smoothing, and an eight-layer rule-based navigation planner. The planner maps perceived hazards to concise commands such as STOP, AVOID, MOVE LEFT/RIGHT, and CONTINUE FORWARD, while a smart text-to-speech gate suppresses passive descriptions when an urgent command is pending. An optional BLIP-based visual validation stage is used in image mode. On the reported 60-frame manually annotated Indian-street evaluation set, the complete pipeline achieves 97.9% hazard recall, 92.8% distance accuracy, 0.407 walkable-corridor IoU, and 84.6 ms average latency per frame. On the IDD detection test set, fine-tuning improves mAP@50 from 0.389 for the COCO- pretrained baseline to 0.542, with mAP@50–95 increasing from 0.224 to 0.351. Ablations show that field-of-view-aware depth calibration is critical for proximity estimation, while temporal tracking is important for instruction stability and motion-aware warnings. The system is deliberately modular: the perception and spatial reasoning stack supplies grounded state information, while deterministic planning converts that state into actionable audio rather than relying on unconstrained generative narration.
研究使用指引
资源信息需要更新?
可报告链接失效、文件异常、信息错误或版本变化,我们会核对并更新资源记录。