Scene Detection
Ask more about what the camera can see.
EchoPath Scene Detection combines on-device visual systems with an optional Qwen3-VL Scene AI path to describe a captured photo and answer follow-up questions.
Whole-scene understanding
From objects to context.
Depending on the selected method and device capability, EchoPath can use Qwen3-VL, Apple Vision, RF-DETR object detection, OCR, colour analysis and image-position information to build a useful description.
The iOS Qwen3-VL model is bundled with EchoPath and is designed to run locally on supported devices. If it cannot complete a request reliably, EchoPath can fall back to its existing vision pipeline.
Search This Photo
Ask follow-up questions about the same image.
The current photo can be explored conversationally instead of requiring a new photo for every question.
Built-in fallback matters.
Large on-device AI models can be demanding. EchoPath’s iOS implementation includes memory-management work, model unloading and fallback behaviour so a useful structured description can remain available on devices where Qwen is not suitable.
Results are estimates.
Object identity, colour, text, materials, position and hazards can all be misunderstood. Scene Detection should be treated as additional information, not guaranteed environmental or safety information.
