Good points! The order is a bit project-dependent though, and where the interest is strongest. From my point of view, Qwen3-1.7B-W4A16 is a bit more interesting as I see a bit more use around it and it can move an actual deploy decision.
But it is not necessary about which accuracy row lands next, but which model x hardware combination we optimize next. Being able to replace an existing deployed model fast with a newer, better one is quite a strong argument, or moving to a cheaper hardware.
In terms of what models may come next, I could see VLAs becoming more interesting going forward since low-latency real-time deployment scenarios are key there.
For example, we recently did a project focusing on VLA (ฯ0.5) optimization on AMD. Some of this was in the AMD keynote last week. We did the GPU VLA inference-latency optimization behind their sub-100ms robot-reasoning number on the Kria AI SOM.
And FYI, missing accuracy benchmarks for the non-LLM/VLM models have been added now. ๐