Settings

Theme

Show HN: VisionLaya: Jev with Vision capabilities

huggingface.co

5 points by someguy101010 · 2 comments · 1 min read

Reader

This model makes calibrated, typed decisions about an image plus optional text. It answers choice, score and noul (yes/no probability) questions in one forward pass, with no text generation.

It adds image input to Laya by replacing Laya's ModernBERT encoder with SmolVLM-256M-Instruct. Laya's predict(state, questions) API, proper-scoring-rule training and temperature calibration are unchanged.

check out the live demo at https://huggingface.co/spaces/thaitea/laya-vision-demo

and the source

- https://huggingface.co/thaitea/laya-vision-smolvlm-256m - https://github.com/r33drichards/laya-vision

1 thread
ranger_danger

This is really quite nice.

Any plans for improvements or training against a larger dataset?

I was hoping to be able to use this for detecting/flagging nsfw image uploads, but I'm still finding too many false positives (that should be obvious to a human) with a low confidence like 0.3 or 0.4.

Thanks!

  • someguy101010OP

    Yea it seems to be very jagged so far but will be training on more datasets.

    I just added evals for existing datasets and some classic controls problems.

    Will likely look for more commercial datasets and synthetic data too.

    I’m funding personally in short term but if I can find some compute can see how well this scales!

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection