Ask HN: Is anyone still working on dLLMs
I remember seeing a lot of hype around Diffusion LLMs and thinking there was a lot of potential there for locally hosting models. Is there some big downside I'm not seeing?
we're full speed ahead building diffusion LLMs at Inception. launched Mercury 2.5 last month and just launched Mercury Voice and Mercury Decide last week. have you tried them?
I tried Mercury 2.5 this week at work for a text extraction / rewriting task. It was happy with the output quality, but the performance wasn't as good as what was claimed, I was able to achieve similar results with Gemini 2.5 FL with much better latency.
thanks for the feedback! would you mind sharing more details about your use case and setup? Mercury 2.5 should be much faster than Gemini 2.5 Flash - AA benchmarks output tokens per second of 617 vs 167). feel free to dm at michael@inceptionlabs.ai and I'll see if I can help.