Types of Extended Reality (XR)
Extended reality (XR) is an umbrella term to refer to augmented reality (AR), virtual reality (VR), and mixed reality (MR).
Virtual Reality (VR) creates a fully immersive digital environment that simulates the real world or an imaginary world. Users wear VR headsets, such as the Oculus Rift or PlayStation VR, which completely block out the physical world and replace it with a virtual one. This technology is used in various fields, including entertainment, healthcare, training, and education. For example, VR can create 360-degree movies, simulate surgeries for medical training, and provide virtual classrooms for remote learning.
Origins of VR: The 1968 3D Head-Mounted Display System created by Sutherland (now dubbed “the godfather of VR”) and an enthusiastic team at Harvard University
AlloSphere (2007)
- 26 stereo projectors with RF glasses
Apple vision pro (2024)
Meta Quest 3 (2023)
PSVR 2 (2023)
and the glass era is coming
Q. In what ways will AR glasses reshape the XR industry and our daily lives?
Components of XR
Vision (HW/SW)
Audio (HW/SW)
Interface (HW/SW)
Algorithm & physics (SW)
- Week 2
Sensory Modalities
Primary senses:
Visual (see), Auditory (hear), and Kinaesthetic (sensations, emotions)
Secondary senses:
Olfactory (smell) and Gustatory (taste).
Vision (HW/SW)
Two eyes, two views → stereo disparity & vergence; roughly 2× geometry/shading without optimizations.
Strict latency budgets (motion‑to‑photon, MTP) → prediction, late latching, reprojection.
Optics in the loop → barrel distortion, CA (Chromatic Aberration) correction.
Comfort constraints → constant frame pacing, low persistence, stable world‑locking.
AR adds reality → camera, depth, SLAM (Simultaneous Localization and Mapping), photometric & geometric registration, occlusion.
Stereoscopic Vision
technique for creating or enhancing the illusion of depth in an image by means of stereopsis for binocular vision
Two eyes capable of facing the same direction to perceive a single three-dimensional image of its surroundings
Stereoscopic rendering in Unity
Strict latency budgets (motion‑to‑photon, MTP)
Strict latency budgets require keeping motion-to-photon (MTP)—from head movement to lit pixels—extremely low, typically under ~20 ms for comfortable VR. This budget spans sensor fusion and pose prediction, app CPU/GPU render time, reprojection (ATW/ASW), display scan-out, and pixel response.
“20 ms all-in” is aggressive but perceptually achievable thanks to prediction, late-latching, and time-warp. In controlled studies, modern HMDs often measure ~20–40 ms motion-to-photon (MTP) at movement onset, with prediction/reprojection reducing the apparent response to single-digit milliseconds shortly after motion begins.
Visuals must be rendered in real time, with low-latency updates to minimize motion sickness.
This requires real-time computation and careful optimization.
Especially for stereoscopic rendering, the computer has to render the scene twice—once per eye.
Trade off
:a situational decision that involves diminishing or losing on quality, quantity, or property of a set or design in return for gains in other aspects
(Latency) Performance vs Rendering quality
(Hardware) Performance vs Mobility
By device (high level):
Quest 2 / Quest 3: Meta relies on late latching and Asynchronous TimeWarp (ATW) to cut perceived MTP; measured MTP on contemporary systems commonly falls in the 20–40 ms band depending on scenario and measurement method. Optimizing VR Graphics with Late Latching | Meta Horizon OS Developers
Apple Vision Pro: Apple’s see-through (camera→display) latency is ~11–12 ms—industry-leading—but note this is photon-to-photon, not full MTP. Independent testing also shows Quest 3 ≈35–40 ms passthrough vs AVP ≈11 ms, and both AVP and Quest 3 score very well on angular MTP for virtual content. Apple Vision Pro Benchmark Test 1: See-Through Latency, Photon-to-Photon | OptoFidelity
Head-pose prediction (sensor fusion), Controller/hand prediction, Late latching & predicted display-time sampling, Late-stage reprojection (timewarp/spacewarp), Eye-gaze prediction for foveated rendering, Network/Cloud-XR prediction...
*** Hitting a true 20 ms end-to-end budget is tough; most headsets land a bit higher.
But with prediction/warp, the perceived latency can approach “instantaneous,” which is why these devices feel responsive in practice.
Optics in the loop → barrel distortion, CA correction
Audio (HW/SW)
Spatial Audio (Stereo and multichannel)
Advantage
No need to wear headset
Disadvantage
Limited spatial resolution
Need many speakers to depict 2D or 3D spatial audio
Hard to personalize the multiuser spatial information
Q. In Unity VR project, were you able to distinguish sound
coming from behind and front?
Were you able to tell if the sound is coming from above or below?
How we are perceiving the sound and spatialize the source:
Left/right uses ITD/ILD (time/level differences).
Front/back & elevation rely on pinna spectral cues + head movement;
(Unity) Project Settings → Audio → Spatializer Plugin (Unity - Manual: Audio spatializers in XR )
Measuring the personal head-related transfer function (HRTF)
Binaural Audio
Using head-related transfer functions (HRTFs), a 3D audio engine filters mono sources with ear-specific timing, level, and spectral cues updated by head pose to recreate convincing front–back and elevation (up/down) localization over headphones.
"Selecting an HRTF spatializer and enabling head tracking lets listeners reliably tell whether a sound is above or below, in front of or behind them."
Interface (HW/SW)
Input/Tracking
6DoF controllers, hand tracking, eye tracking (gaze interaction, foveation), body/face sensors, mic
VIO/SLAM cameras, depth/LiDAR, IMU (accel/gyro)
Feedback (Haptics/Audio)
Haptics (vibration, thermal, force/tethered gloves/suits), HRTF-based 3D audio, spatial voice chat
two black-and-white cameras, two RGB cameras, and a depth sensor
Multiple sensors to capture:
Gesture (Eye, Head, Hand, Body, Feet)
Object
Environment
Main Technology
Sensor Fusion: Combines data from multiple sensors for more accurate detection.
Computer Vision Signal Processing & Machine Learning: Enables cameras data to interpret visual data.
Limitations
Accuracy: Gestural detection can be affected by lighting, occlusion, and sensor limitations.
Latency: Delays between gesture and response can disrupt the sense of immersion.
User Calibration: Systems may require individual user calibration for optimal performance.
Eye Tracking
Gestural Detection in XR
Adrian Bulat and Georgios Tzimiropoulos "Human pose estimation via Convolutional Part Heatmap Regression" (2016)
Back in the days
now the positions of hands are giving
What to do with the position is the different question,
as if data and data processing are different
Processor (HW)
Processor (in headsets): the system-on-chip (SoC) that orchestrates all real-time work—CPU, GPU, and fixed-function blocks—to track, render, and display frames on time.
Quest 2 → Snapdragon XR2 (Gen 1): one-chip mobile SoC handles inside-out tracking and stereo rendering.
Quest 3 / 3S → Snapdragon XR2 Gen 2: much bigger GPU headroom (Qualcomm touts up to 2.5× GPU perf vs XR2 Gen 1) and stronger on-device AI—key for color passthrough MR and hand/scene understanding. (Meta confirms 3S also uses XR2 Gen 2.)
Split compute (M2 + R1): M2 runs apps/graphics/ML; R1 ingests the sensor array and drives displays with ~12 ms photon-to-photon latency—architected for ultra-low-lag passthrough.
Ray-Ban Meta: built on Snapdragon AR1 Gen 1, optimized for low-power capture, voice, and on-glasses AI (no visor-class rendering). Ray-Ban Meta Collection | Qualcomm
Toward thin “true AR” glasses: Snapdragon AR2 Gen 1 uses a distributed, multi-chip design that splits work across the glasses and a host (phone/PC) to cut heat and power.
First Android XR headset from Samsung is slated to use Snapdragon XR2+ Gen 2 (higher clocks, support for up to 4.3K-per-eye @90 Hz, more concurrent cameras)
Qualcomm and Google collaborate to launch new Android XR platform
About this design guidance - Mixed Reality | Microsoft Learn
Still...
Why Quest-like devices don’t use the very latest phone SoCs (System-on-Chip)
Thermals & power on the face. Headsets trap heat; sustained comfort drops as the “micro-climate” under the visor warms, so vendors cap wattage far below phones/laptops. Running cooler beats spiking performance for 30 seconds.
XR-specific pipelines. Chips like Snapdragon XR2 Gen 2 are built for simultaneous camera ingest, low-latency display paths, and reprojection—specialized throughput over peak CPU/GPU clocks. Meta Quest 3 | Qualcomm
Weight, battery, and cost. Pushing a top phone SoC harder would need more cooling and battery on the head (or off-head cabling), hurting comfort and price.
You might be thinking: “Why do we need to know all this hardware? We are software developers.”
Because hardware determines what our software can sense.
Depth sensors give us information about space and geometry.
Hand and gesture tracking capture human actions.
Eye tracking reveals where users are looking.
Spatial audio lets us represent and perceive space through sound.
These technologies give software access to signals from the physical world, the environment, and the user.
So our work should not stop at manipulating Unity scenes and GameObjects.
The real opportunity is to ask:
What can we build when software can sense and respond to the world around us?
This is one of the key promises of XR:
XR is not just about creating virtual worlds.
It can connect the physical and digital worlds.
Examples (Quest3 or 3s):
1. Project Phanto
A mixed reality sample where the user’s real room becomes part of the game environment. Walls, floors, and furniture affect navigation, occlusion, and gameplay.
https://developers.meta.com/horizon/documentation/unity/unity-sample-phanto/
https://github.com/oculus-samples/Unity-Phanto
Provides access to the Quest 3 passthrough camera feed, allowing applications to analyze visual information from the physical environment.
https://developers.meta.com/horizon/documentation/unity/unity-pca-overview/
https://github.com/oculus-samples/Unity-PassthroughCameraApiSamples
Demonstrates how computer vision can detect objects in the physical world and connect the detected objects to positions and content in XR.
https://developers.meta.com/horizon/documentation/unity/unity-sample-camera-object-detection/
A toolkit for working with spatial information about the user’s room, including floors, walls, ceilings, furniture, and other surfaces.
https://developers.meta.com/horizon/documentation/unity/unity-mr-utility-kit-samples/
https://github.com/oculus-samples/Unity-MRUtilityKitSample
Demonstrates how an application can analyze the physical room, identify usable floor areas, and automatically place virtual content in appropriate locations.
https://developers.meta.com/horizon/documentation/unity/unity-sample-mruk-floor-zone/
Uses the geometry and semantic information of the physical room to construct a corresponding virtual environment.
https://developers.meta.com/horizon/documentation/unity/unity-sample-mruk-virtual-home/
Provides environment depth information that allows virtual objects to interact more naturally with real-world geometry, including realistic occlusion.
https://developers.meta.com/horizon/documentation/unity/unity-depthapi-overview/
https://github.com/oculus-samples/Unity-DepthAPI
Turns tracked hand movements into software input, enabling interactions such as grabbing, pinching, pointing, poking, and gesture recognition without controllers.
https://developers.meta.com/horizon/documentation/unity/unity-handtracking-interactions/
https://developers.meta.com/horizon/documentation/unity/unity-isdk-interaction-overview/
https://github.com/oculus-samples/Unity-MoveFast
Allows virtual content to remain associated with specific locations in the physical environment across sessions.
https://developers.meta.com/horizon/documentation/unity/unity-sf-spatial-anchors/
https://github.com/oculus-samples/Unity-StarterSamples
A reference mixed reality project combining scene understanding, passthrough, spatial anchors, and colocation to create experiences shared within the same physical space.
https://github.com/oculus-samples/Unity-Discover
Demonstrates how sound can respond to the listener’s position, head orientation, source location, and acoustic environment to represent space through audio.
https://developers.meta.com/horizon/documentation/unity/unity-sample-meta-xr-audio/
https://github.com/oculus-samples/Unity-MetaXRAudioSDK