With drones, aircraft, vehicle-mounted cameras, 360-degree “action” cameras, and smartphones capturing more video than ever, organizations increasingly want to use video as part of their GIS. Video not only provides a visual, dynamic record of conditions, assets, and events; it captures movement, context, and change over time in ways that individual images cannot. When combined with location information it can support everything from situational awareness and asset inspection to operations and decision-making
ArcGIS works with geospatial video in a number of different ways, and those options are expanding. As the number and types of sensors rapidly grow, and because the dynamic nature of video requires special considerations, it is important to understand the broader geospatial video landscape. Doing so will help you determine the best approach for bringing video into ArcGIS. This blog will discuss key considerations for users in your organization when working with geospatial video, from preparing video for use in ArcGIS to understanding the client apps and capabilities available across the ArcGIS system.
Understanding metadata and choosing the data model
For many organizations, the starting point is to ask: What video do I have, and how can I bring it into ArcGIS?
- Introduction to video and its metadata
- Metadata for geospatial video. To make use of video into a GIS, it needs to include metadata about the location and orientation of the camera. Additionally, the content and accuracy of the metadata is important - the more complete and accurate the metadata, the more functionality users will have in ArcGIS. If the metadata is incomplete, a user may be able to estimate appropriate values, but they will lose some accuracy (and functionality) as a result.
- Managing video and metadata. Because of the dynamic nature of video, there are two data models supported in ArcGIS for managing its metadata, and each has advantages.
- Embedded KLV metadata: The first data model, embedded KLV (Key-Length-Variable), integrates metadata into the video data stream. An important value of embedded metadata is that it can support live streams, such that the sensor location and orientation may be applied (at the client) as the video is received. This data model is also used for stored video files. The KLV format follows a very detailed schema and is originally based on a military specification. Another advantage of this data model is that you don’t need to manage multiple files. A disadvantage is that some video players will not successfully play the video file/stream, and the schema is not optimized for some video modes (e.g. 360 video).
- External metadata: The second data model is to manage and access metadata external to the video file. This is typically the case for your raw data (a simple video file plus its metadata, often in *.csv format). This raw data can be processed into the embedded KLV data model or managed separately from the video via a feature class using the ArcGIS oriented imagery capability. An advantage of this method is greater simplicity and more flexible support for video that can’t be mapped onto the ground - for example, 360 video and oblique views aimed above the horizon such as from mobile dashcams or asset inspections.
- Preparing your video data for ArcGIS
This section provides guidance on choosing which data model for metadata+video best suits the data. This advice is based primarily on the collection platform (aerial vs. terrestrial), the view orientation of the video, and available metadata. Based on these recommendations, users can then consider which of the client applications and features available within ArcGIS best satisfy their requirements.
- Aerial platforms. The embedded KLV metadata is often the best choice for video collected from an aircraft or drone, although it can also be managed with external metadata using the oriented imagery capability.
Professional and military-grade systems often encode KLV metadata directly into video files or livestreams when the video is captured. Client applications that work with embedded KLV format (Excalibur and ArcGIS Pro with Image Analyst) can use this video immediately, without any additional processing.
Drone video captured with ArcGIS Flight is automatically accompanied by a geospatial video log that supports either data model. The Video Multiplexer can be used to embed the metadata as KLV, or Add Images to Oriented Imagery Dataset (with the AerialFrameVideo category) can access the video using oriented imagery.
For video and metadata from another airborne system, users will need to format the available metadata to match the supported aerial schema. Doing so preserves the option to use either data model.
- Terrestrial platforms. Video from moving terrestrial (ground level) platforms, such as vehicle-mounted or 360-degree cameras, is typically best utilized in ArcGIS when it is managed with external metadata in an oriented imagery dataset. These videos can be added using the Add Images to Oriented Imagery Dataset, selecting either the TerrestrialFrameVideo or Terrestrial360Video category as appropriate.
For video captured from a fixed terrestrial location, Generate Video Metadata can create a metadata table formatted according to the aerial schema. External metadata in oriented imagery is generally recommended for this type of video because these cameras often include view orientations above the horizon.
- View orientation above/below horizon. In addition to the collection platform, an important consideration is the primary view direction of the video. If the camera is aimed mostly toward the ground, either data model can be used. However, if the video frequently aims toward the horizon, oriented imagery is likely to be preferred. This is because many features implemented in ArcGIS for embedded KLV video assume that the video footprint is projected onto the ground.
- Metadata content/completeness. The functionality available in ArcGIS depends largely on the completeness of the metadata that accompanies the video. Depending on your video hardware, capturing metadata that is complete, accurate, and in an easily usable format is one of the key challenges for many video systems.
- Complete metadata. ArcGIS uses a common set of metadata fields to describe the location, orientation, and viewing characteristics of a video sensor. It is recommended that a complete metadata record should follow the aerial platform schema (recommended even if the video was not captured from an aerial platform, since this schema addresses many important parameters for video as a geospatial data type). Metadata will include:
- Camera location (x,y,z) and timestamp
- Platform orientation
- Sensor orientation relative to the platform
- Camera field of view
Some values may be fixed or estimated – common examples include sensor orientation relative to the platform (if there is no moving gimbal), or the camera field of view (if no zoom lens). More complete and accurate values improve the placement of the sensor, viewing direction, and video footprint.
- Limited metadata. Many terrestrial systems record only camera location and time, commonly in comma-separated value (CSV) or GPS exchange format (GPX) files. This metadata can be used in an oriented imagery dataset when added with the geoprocessing tool under the TerrestrialFrameVideo or Terrestrial360Video category as appropriate.
Some systems also report heading, and additional values can be estimated. For example, a fixed heading offset may be added for a side-facing vehicle camera, or a fixed pitch angle for a downward-facing camera. Heading values for the two Terrestrial categories can be estimated by ArcGIS during the “Add Images” process.
- Minimum metadata. A common question is, “What is the minimum metadata required?” but there is no single answer for minimum metadata content. ArcGIS Video Server can manage and share video with no metadata, and it can be viewed in ArcGIS Pro. For geospatial applications, the practical minimum is camera location and timestamp. Additional metadata progressively enables more functionality:
- (x,y, and optionally z) location updates the sensor position on the map.
- Heading and other orientation angles identify the viewing direction.
- Field of view supports an estimated view footprint.
Complete orientation information improves map-to-video positioning.
Comparing features of the client apps in ArcGIS for video managed with embedded KLV vs. oriented Imagery
The numerous client apps in ArcGIS for geospatial video provide different capabilities, and their features depend on the data model and metadata available with the video. The table below summarizes some of the key features and differences, referenced to the client applications across ArcGIS.
For either data model,
- The usable extents of the video can be shown on the map;
- The video plays in a separate window, with a moving indicator on the map for camera location;
- Overlay of GIS features and map-to-video coordinate identification require sufficient camera position, orientation, and field-of-view metadata.
Continue with the table below for additional detail about software features.
Notes:
- For sharing video via the web using oriented imagery (see asterisks * above), the video must be web-accessible and shared publicly (until November 2026), and the oriented imagery dataset must be published as a layer before it can be used in supported clients. After November 2026, oriented imagery will support secure data storage via ArcGIS Enterprise.
- For embedded KLV videos, ArcGIS Pro requires the Image Analyst extension (the “Motion Imagery” tools).
- ArcGIS Video Server does not appear as a separate feature column because it is a hosting and streaming infrastructure, not an interactive client. As of version 12.2 (late 2026), Video Server does not support the oriented imagery data model for video, but this will likely change in future versions.
- Scene Viewer via Portal (ArcGIS Online or Enterprise) is not shown for oriented imagery since it did not support video in oriented imagery until the November 2026 release of ArcGIS Online. Release version for support in Enterprise is TBD.