Settings

Theme

Show HN: Stop GPU pods placement getting bottlenecked by reserved VRAM

2 points by medicis123 4 months ago · 0 comments · 1 min read


We have built a GPU Runtime for Nvidia GPUs that can run multiple development/experimental/inference workloads per GPU with safe overcommit of VRAM, dynamic fractional allocation of GPU cores, and Deduplication of weights in VRAM.

We are looking for teams to give it a try.

More details to get a trial license - https://www.woolyai.com.

No comments yet.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection