btw this is a successor to v1 from about a year ago which used gpt4 to generate primitives and arranged them in a scene.
however, i realized regardless of how good of a language model you use it won't quite beat a fully 3D diffusion model. hence the v2, now much more capable.