InstantXR: Instant XR Environment on the Web Using Hybrid Rendering of Cloud-based NeRF with 3D Assets | Proceedings of the 27th International Conference on 3D Web Technology

· ACM Conferences

6 min read Original article ↗

Abstract

Abstract

For an XR environment to be used on a real-life task, it is crucial all the contents are created and delivered when we want, where we want, and most importantly, on time. To deliver an XR environment faster and correctly, the time spent on modeling should be considerably reduced or eliminated. In this paper, we propose a hybrid method that fuses the conventional method of rendering 3D assets with the Neural Radiance Fields (NeRF) technology, which uses photographs to create and display an instantly generated XR environment in real-time, without a modeling process. While NeRF can generate a relatively realistic space without human supervision, it has disadvantages owing to its high computational complexity. We propose a cloud-based distributed acceleration architecture to reduce computational latency. Furthermore, we implemented an XR streaming structure that can process the input from an XR device in real-time. Consequently, our proposed hybrid method for real-time XR generation using NeRF and 3D graphics is available for lightweight mobile XR clients, such as untethered HMDs. The proposed technology makes it possible to quickly virtualize one location and deliver it to another remote location, thus making virtual sightseeing and remote collaboration more accessible to the public. The implementation of our proposed architecture along with the demo video is available at https://moonsikpark.github.io/instantxr/.

AI Summary

AI-Generated Summary (Experimental)

This summary was generated using automated tools and was not authored or reviewed by the article's author(s). It is provided to support discovery, help readers assess relevance, and assist readers from adjacent research areas in understanding the work. It is intended to complement the author-supplied abstract, which remains the primary summary of the paper. The full article remains the authoritative version of record. Click here to learn more.

Click here to comment on the accuracy, clarity, and usefulness of this summary. Doing so will help inform refinements and future regenerated versions.

To view this AI-generated plain language summary, you must have Premium access.

Formats available

You can view the full content in the following formats:

References

[1]

Volga Aksoy and Dean Beeler. 2015. INTRODUCING ASW 2.0: BETTER ACCURACY, LOWER LATENCY. Retrieved September 29, 2022 from https://www.oculus.com/blog/introducing-asw-2-point-0-better-accuracy-lower-latency/

[2]

Michael Antonov. 2015. Asynchronous Timewarp Examined. Retrieved September 29, 2022 from https://developer.oculus.com/blog/asynchronous-timewarp-examined/

[3]

Lucio Azzari, Federica Battisti, and Atanas Gotchev. 2010. Comparative Analysis of Occlusion-Filling Techniques in Depth Image-Based Rendering for 3D Videos. In Proceedings of the 3rd Workshop on Mobile Video Delivery (Firenze, Italy) (MoViD ’10). Association for Computing Machinery, New York, NY, USA, 57–62. https://doi.org/10.1145/1878022.1878037

[4]

Ricardo Cabello. 2022. Three.js. Retrieved July 30, 2022 from https://github.com/mrdoob/three.js/

[5]

Chun-Min Chang. 2022. Firefox Bugzilla: [meta] Tracking bug for WebCodecs API implementation. Retrieved September 29, 2022 from https://bugzilla.mozilla.org/show_bug.cgi?id=WebCodecs

[6]

F Condorelli, F Rinaudo, F Salvadore, and S Tagliaventi. 2021. A Comparison Between 3d Reconstruction Using Nerf Neural Networks and Mvs Algorithms on Cultural Heritage Images. The International Archives of Photogrammetry, Remote Sensing and Spatial Information Sciences 43(2021), 565–570. https://doi.org/10.5194/isprs-archives-XLIII-B2-2021-565-2021

[7]

Chris Cunningham, Paul Adenot, and Bernard Aboba. 2022. WebCodecs. Retrieved July 31, 2022 from https://www.w3.org/TR/webcodecs/

[8]

Chris Cunningham and Dan Sanders. 2022. Chrome Platform Status Feature: WebCodecs. Retrieved September 29, 2022 from https://chromestatus.com/feature/5669293909868544

[9]

Brian Curless and Marc Levoy. 1996. A Volumetric Method for Building Complex Models from Range Images. In Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques(SIGGRAPH ’96). Association for Computing Machinery, New York, NY, USA, 303–312. https://doi.org/10.1145/237170.237269

[10]

Stephan J. Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien Valentin. 2021. FastNeRF: High-Fidelity Neural Rendering at 200FPS. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 14346–14355.

[11]

Cullen Jennings, Henrik Boström, and Jan-Ivar Bruaroey. 2022. WebRTC 1.0: Real-Time Communication Between Browsers. Retrieved July 30, 2022 from https://www.w3.org/TR/webrtc/

[12]

Brandon Jones, Manish Goregaokar, and Rik Cabanier. 2022. WebXR Device API. Retrieved July 30, 2022 from https://www.w3.org/TR/webxr/

[13]

Yongjae Lee and Byounghyun Yoo. 2021. XR collaboration beyond virtual reality: work in the real world. Journal of Computational Design and Engineering 8, 2(2021), 756–772. https://doi.org/10.1093/jcde/qwab012

[14]

Yongjae Lee, Byounghyun Yoo, and Soo-Hong Lee. 2021. Sharing Ambient Objects Using Real-Time Point Cloud Streaming in Web-Based XR Remote Collaboration. In The 26th International Conference on 3D Web Technology (Pisa, Italy) (Web3D ’21). Association for Computing Machinery, New York, NY, USA, Article 4, 9 pages. https://doi.org/10.1145/3485444.3487642

[15]

Diego Marcos, Don McCurdy, and Kevin Ngo. 2022. A-Frame. Retrieved July 30, 2022 from https://github.com/aframevr/aframe

[16]

Microsoft. 2022. Babylon.js. Retrieved July 30, 2022 from https://github.com/BabylonJS/Babylon.js/

[17]

Ben Mildenhall, Peter Hedman, Ricardo Martin-Brualla, Pratul P. Srinivasan, and Jonathan T. Barron. 2022. NeRF in the Dark: High Dynamic Range View Synthesis From Noisy Raw Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 16190–16199. https://openaccess.thecvf.com/content/CVPR2022/papers/Mildenhall_NeRF_in_the_Dark_High_Dynamic_Range_View_Synthesis_From_CVPR_2022_paper.pdf

[18]

Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2020. Nerf: Representing scenes as neural radiance fields for view synthesis. In Computer Vision – ECCV 2020. Springer International Publishing, Cham, 405–421. https://doi.org/10.1007/978-3-030-58452-8_24

[19]

MPEG. 2019. Information technology — Dynamic adaptive streaming over HTTP (DASH) — Part 1: Media presentation description and segment formats. Standard ISO/IEC 23009-1:2019. International Organization for Standardization, Geneva, CH. https://www.iso.org/standard/79329.html

[20]

Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics 41, 4 (jul 2022), 1–15. https://doi.org/10.1145/3528223.3530127

[21]

Roger Pantos and William May. 2017. HTTP Live Streaming. RFC 8216. https://doi.org/10.17487/RFC8216

[22]

Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M. Seitz. 2021. HyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance Fields. ACM Trans. Graph. 40, 6, Article 238 (dec 2021). https://doi.org/10.48550/arXiv.2106.13228

[23]

Anup Rao, Rob Lanphier, and Henning Schulzrinne. 1998. Real Time Streaming Protocol (RTSP). RFC 2326. https://doi.org/10.17487/RFC2326

[24]

Michael Schmeing. 2011. Depth Image Based Rendering. 279–310. https://doi.org/10.1007/978-3-642-22407-2_12

[25]

Johannes Lutz Schönberger and Jan-Michael Frahm. 2016. Structure-from-Motion Revisited. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4104–4113. https://doi.org/10.1109/CVPR.2016.445

[26]

Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. 2016. Pixelwise View Selection for Unstructured Multi-View Stereo. In European Conference on Computer Vision (ECCV). https://doi.org/10.1007/978-3-319-46487-9_31

[27]

Maria Sharabayko, Maxim Sharabayko, Jean Dube, Joonwoong Kim, and Jeongseok Kim. 2021. The SRT Protocol. Internet-Draft draft-sharabayko-srt-01. Internet Engineering Task Force. https://datatracker.ietf.org/doc/draft-sharabayko-srt/01/ Work in Progress.

[28]

Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P. Srinivasan, Jonathan T. Barron, and Henrik Kretzschmar. 2022. Block-NeRF: Scalable Large Scene Neural View Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8248–8258. https://openaccess.thecvf.com/content/CVPR2022/papers/Tancik_Block-NeRF_Scalable_Large_Scene_Neural_View_Synthesis_CVPR_2022_paper.pdf

[29]

A. Tewari, J. Thies, B. Mildenhall, P. Srinivasan, E. Tretschk, W. Yifan, C. Lassner, V. Sitzmann, R. Martin-Brualla, S. Lombardi, T. Simon, C. Theobalt, M. Nießner, J. T. Barron, G. Wetzstein, M. Zollhöfer, and V. Golyanik. 2022. Advances in Neural Rendering. Computer Graphics Forum 41, 2 (2022), 703–735. https://doi.org/10.1111/cgf.14507 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/cgf.14507

[30]

J. M. P. van Waveren. 2016. The Asynchronous Time Warp for Virtual Reality on Consumer Hardware. In Proceedings of the 22nd ACM Conference on Virtual Reality Software and Technology (Munich, Germany) (VRST ’16). Association for Computing Machinery, New York, NY, USA, 37–46. https://doi.org/10.1145/2993369.2993375

[31]

WHATWG. 2022. WebSockets. Retrieved July 30, 2022 from https://websockets.spec.whatwg.org/

[32]

Yingen Xiong and Christopher Peri. 2021. Space-Warp with Depth Propagation in XR Applications. In 2021 IEEE International Symposium on Multimedia (ISM). 50–57. https://doi.org/10.1109/ISM52913.2021.00017

[33]

Xiaokun Xu and Mark Claypool. 2021. A First Look at the Network Turbulence for Google Stadia Cloud-based Game Streaming. In IEEE INFOCOM 2021 - IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). 1–5. https://doi.org/10.1109/INFOCOMWKSHPS51825.2021.9484481

[34]

Minxia Yang, Jiaqi Zhang, and Lu Yu. 2019. Perceptual Tolerance to Motion-To-Photon Latency with Head Movement in Virtual Reality. In 2019 Picture Coding Symposium (PCS). 1–5. https://doi.org/10.1109/PCS48520.2019.8954518