Settings

Theme

Show HN: I wrote a 1-bit WebGPU runtime to run a 1.7B LLM in the browser

aidekin.com

5 points by stfurkan a month ago · 6 comments

Reader

debdattabasu a month ago

Wow this is a cool prototype and looks amazing. I am surprised some level of baseline intelligence survives this kind of aggro quantization. Do you think 300mb initial download is ok for something like quick website where I want to ask quick support question? Are you planning to have hosted fallback to answer q while download is happening?

  • stfurkanOP a month ago

    Thank you :) Currently I am not planning to have hosted fallback Q&A but it's a nice idea. It should only download the ~300mb initially once and then use the cached model.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection