TIE first sets a new state of the art on standard single-turn CIR benchmarks, reaching 39.4 mAP@5 on CIRCO and 38.7 R@1 on CIRR, proving its representations are strong even before any multi-turn context comes into play.
Single-turn CIR across three benchmarks. TIE is the most consistent method, leading on CIRR and CIRCO and staying competitive on FIQ, whereas prior methods excel on one benchmark but fall behind on others.
The real story, however, is multi-turn. On the new CoCo-IR benchmark, prior CIR models must be adapted with proxy inputs, and they collapse as the dialogue deepens, even when given an oracle Gemini-2.5-Pro summary of the full history. TIE instead ingests the entire context natively and stays robust, keeping the gap wide all the way to the fourth turn.
Recall@1 across turns. The best baseline is shown under three adaptation strategies; all degrade sharply with depth, while TIE’s native full-context modeling holds up. Multi-turn visual dialogue cannot be losslessly compressed into a single text prompt.