Scene Detection

Ask more about what the camera can see.

EchoPath Scene Detection combines on-device visual systems with an optional Qwen3-VL Scene AI path to describe a captured photo and answer follow-up questions.

Whole-scene understanding

From objects to context.

Depending on the selected method and device capability, EchoPath can use Qwen3-VL, Apple Vision, RF-DETR object detection, OCR, colour analysis and image-position information to build a useful description.

On-device Qwen model

The iOS Qwen3-VL model is bundled with EchoPath and is designed to run locally on supported devices. If it cannot complete a request reliably, EchoPath can fall back to its existing vision pipeline.

Search This Photo

Ask follow-up questions about the same image.

The current photo can be explored conversationally instead of requiring a new photo for every question.

What is at the bottom right?What does the sign say?How many chairs can you see?What colour is the bag?What is behind the chair?Describe everything on the table.

Built-in fallback matters.

Large on-device AI models can be demanding. EchoPath’s iOS implementation includes memory-management work, model unloading and fallback behaviour so a useful structured description can remain available on devices where Qwen is not suitable.

Results are estimates.

Object identity, colour, text, materials, position and hazards can all be misunderstood. Scene Detection should be treated as additional information, not guaranteed environmental or safety information.