Settings

Theme

Show HN: WavexAI – Unlimited tokens for a fixed cost

wavexai.dev

6 points by wilsprouse · 12 comments

Reader

5 threads
wilsprouseOP

I have never understood the per token pricing model, so I built WavexAI to see what I was missing. Maybe I'll find something at scale, but turns out you can provide fixed cost inference without the fear of running out of tokens.

Would love feedback, and testers on the site!

stackfrost

You are yourself in some pay paying for the compute right? Isn't this fix price an opportunity for many to abuse?

How are you managing rate limits and what model are you using? if you would like to disclose that.

  • wilsprouseOP

    Yeah, of course we are paying for the compute on our end. I think we are waiting for our users to stress test our theory, however our secret sauce is optimizing inference so that this is a feasible product on both ends.

    Like I alluded to in my initial comment, I've never understood why the inference providers of today charge per token. Whichever marketing department got our industry to be ok with being charged for the output of a REST call, I applaud.

    I'm not sure it will all hold up at scale, but I am still waiting for the test to break. It helps to be a curious and optimistic person in this endeavor.

    Not currently implementing any rate limits. Using an assortment of open models.

prologic

What's your cost model / pricing based on here? If you're going to provide $20/month for unlimited APi calls, I can see this getting abused very reasily.

  • wilsprouseOP

    Pricing is largely based on how many users we think we can fit per our baseline modset, and how much the cost of that modset is determines where we can work back from in price. From there, we would scale on a per modset basis.

    It may not turn out to be a perfect science, but in the name of shipping something and getting feedback, it's out there

    • prologic

      Sorry, but what's a "modset"? I _assume_ you're renting GPU(s) from a IaaS provider?

      • wilsprouseOP

        Sorry, that term may be specific to the industry I work in haha. I mean it as the "modification Set", or configuration set rather. Basically just what's our baseline configuration in terms of the entire software/hardware stack.

        Yeah right now we are renting GPUs, but at scale I would think we would own our own GPUs. As with any business, would probably go with the cheapest viable option

redditgrowthhub

How does it work and how are you getting more people to discover it

  • wilsprouseOP

    It is very new (launched a couple days ago), so still trying to figure out how to get it out there. I'm a software nerd not a marketing genius unfortunately.

    As for how it works, its a pretty standard inference provider, where we offer a chatbot and an API, at unlimited usage for a fixed cost. Since its very new, we're still wondering if there's something we're missing as for why fixed cost AI is not the norm.

    Hoping it is similar to the music industry, where you once had to pay $1.99 to download a song, and now you can stream all you want for $12.99. We want to begin that new trend.

weedfroglozenge

What models are available for use?

  • wilsprouseOP

    We kinda take the stance that the model is a commodity, and could be swapped at any point so we don't really advertise what model we are currently using but its no secret so since you asked we are using the Qwen family models.

    I don't love using a model from china for many reasons (security, software supply chain), but unfortunately thats the state of the open source model landscape.

    Also in the future we'd like to offer multiple models

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection