Hosting only the best open source model — DefaultModel

DefaultModel

1 min read Original article ↗

Open source models are amazing but hosting them reliably is not. Closed source labs are reliable and low cost as they can aggregate volume mostly to their latest model. At the pace of open source development, current inference providers constantly have to host multiple models, optimize for all of them, and split their GPU resources to serve them all. We're taking a different approach. We'll only host one model based, the best open source model voted by the community, and optimize specifically for it while providing all the GPU resources to it giving you reliability of a large lab for open models.

->

Zero data retention by design

->

99.9% uptime with no degraded serving

Production performance, all the time

One model means all our capacity serves it: 99.9% uptime, high token throughput in and out, and no degraded or quantized serving.