Overview
Media toolkit for processing video and audio attachments.
This package provides the infrastructure for media processing, including video clip extraction, frame sampling, audio extraction, keyframe extraction, and temporal segmentation.
Submodules
processor-- Processor families for video clip, frame sampling, audio extraction, deinterlace, frame extraction, and dense frame decode.segmenter-- Temporal segmenters for splitting media into fixed-duration or shot-based chunks.keyframe_extractor-- Keyframe extraction from video streams.
Usage
from gllm_multimodal.media_toolkit import processor, segmenter
AudioExtractionProcessor()
Bases: BackendSelectableProcessor[Attachment, Attachment], ABC
Family base for extracting audio tracks from video attachments.
This class serves as a unified entry point for audio extraction operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.
Why use this base class?
- Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
- Simplicity: No need to handle fallback logic or conditional imports yourself.
- Future-proofing: New backends can be added to the library without requiring changes to your application code.
Usage Example
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment
# Instantiates the best available backend automatically
processor = AudioExtractionProcessor.build()
attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment
# Explicitly force the ffmpeg backend
processor = AudioExtractionProcessor.build(backend="ffmpeg")
attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment
# Explicitly force the moviepy backend
processor = AudioExtractionProcessor.build(backend="moviepy")
attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
BaseSegmenterConfig
Bases: BaseModel
Empty base config used to carry values into a concrete segmenter config.
Concrete segmenter configs declare and validate only the settings their
implementation consumes. Extra fields are retained here so callers can
up-cast an exact base config with from_dict.
from_dict(data=None)
classmethod
Build a validated config from a dict, base config, or existing instance.
Sibling subclass configs are rejected: re-validating an unrelated
model would silently drop (or carry over) fields the caller never
set. Pass a dict or an exact BaseSegmenterConfig to up-cast
into the concrete subclass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Self | BaseSegmenterConfig | dict[str, Any] | None
|
Raw config
payload. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
Self |
Self
|
Validated config of the concrete subclass. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
Detector
Bases: StrEnum
Supported shot-detector algorithm names.
FixedDurationSegmenter(config=None)
Bases: BaseSegmenter[FixedDurationSegmenterConfig]
Segment attachments using explicit per-segment durations.
segment
returns cumulative time windows from config.segment_durations.
materialize
clips each window into a separate attachment using the cached
VideoClipProcessor obtained from CompositeMediaMixin.
config_model()
classmethod
Return this segmenter's concrete config model.
Returns:
| Type | Description |
|---|---|
type[FixedDurationSegmenterConfig]
|
type[FixedDurationSegmenterConfig]: The model used to validate fixed-duration segmenter configuration. |
materialize(attachment, segment, segment_index=0)
async
Clip and return one attachment for a precomputed segment window.
This method is convenient when segment planning and clip extraction are performed in separate stages, and only selected windows should be materialized.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source media attachment to clip. |
required |
segment
|
VideoSegment
|
Segment boundary plan to materialize. |
required |
segment_index
|
int
|
Zero-based index used in generated output filenames. Defaults to 0. |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Clipped attachment with
[ |
process(attachment, **kwargs)
async
Materialize fixed-duration clips from one media attachment.
Unlike calling
segment
directly, this method returns real clipped attachment outputs with
[VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment] metadata embedded on
each result.
It is the main runtime entrypoint when you need files/bytes for every
configured duration window, not only boundary plans.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source media attachment to split. |
required |
**kwargs
|
Any
|
Forwarded processing arguments accepted by the base media-toolkit contract. |
{}
|
Notes
- Delegates shared validation and orchestration to
MediaToolkit.process(inherited byBaseSegmenter).
Returns:
| Type | Description |
|---|---|
list[Attachment]
|
list[Attachment]: One clipped attachment per configured segment window. |
segment(attachment)
async
Return computed fixed windows without creating clip attachments.
This is useful for previewing time boundaries (for inspection, logging, or downstream planning) before paying the cost of media clipping.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source media attachment. The payload itself is not read by this implementation when computing boundaries. |
required |
Notes
- Delegates shared validation to
BaseSegmenter.segment.
Returns:
| Type | Description |
|---|---|
list[VideoSegment]
|
list[VideoSegment]: Fixed cumulative windows derived from |
list[VideoSegment]
|
|
FixedDurationSegmenterConfig
Bases: BaseSegmenterConfig
Configuration for FixedDurationSegmenter.
Requires a non-empty segment_durations list.
Attributes:
| Name | Type | Description |
|---|---|---|
segment_durations |
list[float]
|
Ordered segment durations in seconds. |
start_time |
float
|
Base start time for the first segment. Defaults to 0.0. |
validate_segment_durations(value)
classmethod
Reject non-positive segment durations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
list[float]
|
Candidate segment durations. |
required |
Returns:
| Type | Description |
|---|---|
list[float]
|
list[float]: The durations unchanged when valid. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If any duration is non-positive. |
FrameDecodeProcessor()
Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC
Family base for dense video-frame decoding.
Why use this base class?
- Portability: Swap FFmpeg vs GStreamer without changing call sites.
- FIPS choice:
backend="gstreamer"decodes with system plugins (no bundled-FFmpeg wheels);backend="ffmpeg"shells out to the systemffmpegbinary. - Separation: Frame decoding stays here; frame scoring (shot detection, keyframes) lives in segmenters/extractors.
Usage
from gllm_multimodal.media_toolkit.processor.frame_decode_processor import (
FrameDecodeProcessor,
)
processor = FrameDecodeProcessor.build(backend="gstreamer")
frames = await processor.process(video_attachment)
frame_metadata(frame_index, fps)
Build the metadata dict attached to every decoded frame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
frame_index
|
int
|
Zero-based position in decode order. |
required |
fps
|
float | None
|
Effective sampling rate, if known. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
dict[str, Any]: |
iter_rgb_frames_sync(attachment, *, sample_fps=None, target_width=None, as_rgb=True)
Yield decoded frames after the backend writes the full PNG sequence.
The decoder subprocess/pipeline completes first, so peak temp-disk
usage is every sampled PNG at once. After that, this generator opens
each file in order, yields the payload, and deletes the PNG so Python
RAM stays O(1) in frames. Prefer this over process when callers
only need a scored stream and can tolerate the peak-disk cost.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
sample_fps
|
int | None
|
Override downsample rate. Defaults to the constructor config. |
None
|
target_width
|
int | None
|
Override downscale width. Defaults to the constructor config. |
None
|
as_rgb
|
bool
|
When True (default), yield RGB arrays.
When False, yield raw PNG bytes (used by |
True
|
Yields:
| Type | Description |
|---|---|
tuple[Any, dict[str, Any]]
|
tuple[Any, dict[str, Any]]: |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If decoding fails or yields no frames. |
FileNotFoundError
|
If a required decoder binary is missing. |
FrameExtractionProcessor()
Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC
Family base for extracting image frames at timestamps.
Why use this base class?
- Portability: Swap FFmpeg vs GStreamer without changing call sites.
- Batch I/O: Extract many keyframes in one
processcall. - Separation: Keyframe planning stays in extractors; decode is here.
Usage
from gllm_multimodal.media_toolkit.processor.frame_extraction_processor import (
FrameExtractionProcessConfig,
FrameExtractionProcessor,
)
processor = FrameExtractionProcessor.build()
frames = await processor.process(
video_attachment,
process_config=FrameExtractionProcessConfig(timestamps=[1.5, 4.0]),
)
set_timestamps(timestamps)
abstractmethod
Configure default timestamps used when process_config is omitted.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamps
|
list[float]
|
Non-empty list of non-negative seconds. |
required |
stream_frames(attachment, sample_fps, deinterlace=None)
abstractmethod
Stream frames sampled at a fixed rate from a single decode pass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
sample_fps
|
float
|
Positive sampling rate in frames per second. |
required |
deinterlace
|
DeinterlaceMode | None
|
Deinterlace override. Defaults to None. |
None
|
Returns:
| Type | Description |
|---|---|
AsyncIterator[tuple[float, ndarray]]
|
AsyncIterator[tuple[float, np.ndarray]]: Timestamps and BGR |
FrameSamplingProcessor()
Bases: BackendSelectableProcessor[Attachment, Attachment], ABC
Family base for resampling video attachments to a target frame rate.
This class serves as a unified entry point for frame sampling operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.
Why use this base class?
- Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
- Simplicity: No need to handle fallback logic or conditional imports yourself.
- Future-proofing: New backends can be added to the library without requiring changes to your application code.
Usage Example
from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import (
FrameSamplingProcessor,
GstFrameSamplingConfig,
)
from gllm_inference.schema import Attachment
# Instantiates the best available backend automatically
processor = FrameSamplingProcessor.build(
config=GstFrameSamplingConfig(default_target_fps=2)
)
attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import FrameSamplingProcessor
from gllm_inference.schema import Attachment
# Explicitly force the ffmpeg backend
processor = FrameSamplingProcessor.build(backend="ffmpeg")
attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)
MediaToolkit()
Bases: ABC, Generic[T_in, T_out]
Base abstraction for all media toolkit processing components.
This class provides the shared lifecycle and registry behavior used by both: - concrete leaf processors (e.g. backend-specific audio/video processors), and - composite components (e.g. segmenters, keyframe extractors) that orchestrate nested processors.
Key responsibilities:
- auto-register subclasses by class name for class-name-based construction via
build;
- provide consistent input validation against supported_mimetypes;
- define async processing contracts through process and
process_batch.
Contributor guidance:
- inherit this class directly for concrete processors with custom behavior;
- inherit BackendSelectableProcessor when one logical processor family maps
to multiple backend implementations;
- inherit composite bases (e.g. BaseSegmenter) for orchestration-style
components.
Example
Building a processor by class name
from gllm_multimodal.media_toolkit.media_toolkit import MediaToolkit
processor = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
result = await processor.process(video_attachment)
Checking mimetype support
if processor.is_supported(attachment):
result = await processor.process(attachment)
Listing registered processors
print(list(MediaToolkit.registry.keys()))
# ['GstAudioExtractionProcessor', 'GstVideoClipProcessor', ...]
Initialize processor logging.
name
property
Return a stable component name for optional plan metadata.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Component class name. |
registry = {}
class-attribute
Global class-name registry used by build.
Maps each registered subclass name to its concrete class type, enabling
string-based construction such as MediaToolkit.build("AudioExtractionProcessor").
supported_mimetypes = ['*/*']
class-attribute
MIME types this processor accepts (supports wildcards, e.g. 'video/*').
Defaults to ['*/*'] (accept all). Override as a class attribute in subclasses.
__init_subclass__(**kwargs)
Register every concrete subclass into the global registry.
This hook is triggered automatically by Python whenever a class inherits
from MediaToolkit (directly or indirectly). Registration happens at
class definition/import time, so classes become immediately discoverable
by build
without manual setup.
Registration key
- The subclass'
__name__(e.g."AudioExtractionProcessor"). - The value stored is the subclass type itself.
Why uniqueness is enforced
build(class_name=...)uses this registry for class resolution.- Duplicate class names would silently shadow earlier classes and could route builds to unintended implementations.
- To prevent that ambiguity, duplicate keys raise
TypeError.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
**kwargs
|
Any
|
Extra class declaration keyword arguments forwarded to
parent |
{}
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If a subclass with the same class name is already registered, preventing silent dispatch to the wrong implementation. |
available_backends_for(class_name)
classmethod
Return backend keys registered for a processor family class name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
class_name
|
str
|
Registered processor family class name. |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: Available backend keys. Empty when the class is unknown or not a backend-selectable family base. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the class name is unknown. |
Example
backends = MediaToolkit.available_backends_for("VideoClipProcessor")
print(backends)
["gstreamer", "ffmpeg"]
build(class_name, backend=None, **kwargs)
classmethod
Build a processor by class name.
Family abstract classes (e.g. AudioExtractionProcessor) resolve a concrete
backend implementation via backend. Composite components (segmenters,
keyframe extractors) store backend on the instance for nested resolution.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
class_name
|
str
|
Registered subclass name. |
required |
backend
|
str | MediaBackend | None
|
Backend key for family classes, or nested processor preference for composite instances. Defaults to None. |
None
|
**kwargs
|
Any
|
Constructor kwargs passed to the processor class. |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
MediaToolkit |
MediaToolkit
|
Instantiated processor. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the class name is unknown. |
Example
proc = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
seg = MediaToolkit.build(
"FixedDurationSegmenter",
backend="gstreamer",
config={"segment_durations": [2.0]},
)
print(proc)
print(seg)
<GstAudioExtractionProcessor instance>
<FixedDurationSegmenter instance>
build_from_registry(backend=None, **kwargs)
classmethod
Instantiate this registered class.
Subclasses override this hook to customize registry-based construction
(e.g. backend-selectable families resolve a concrete backend; composites
store backend for nested processor resolution).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
backend
|
str | MediaBackend | None
|
Backend key forwarded to subclass overrides. Ignored by the base implementation. |
None
|
**kwargs
|
Any
|
Constructor kwargs passed to the processor class. |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
MediaToolkit |
MediaToolkit
|
Instantiated processor. |
Example
# Called indirectly by MediaToolkit.build(...)
processor = SomeRegisteredProcessor.build_from_registry(custom_flag=True)
print(processor)
<SomeRegisteredProcessor instance>
is_supported(attachment)
Return whether the attachment's mimetype is accepted by this processor.
Callers can use this to check compatibility before calling
process or process_batch, avoiding a ValueError.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
The attachment to check. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
|
Example
ok = processor.is_supported(Attachment(mime_type="video/mp4"))
print(ok)
True
list_available_backends()
classmethod
Return backend keys when this class supports backend selection.
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: Available backend keys. Empty for classes that are not backend-selectable family bases. |
Example
Base classes are not backend-selectable.
print(MediaToolkit.list_available_backends())
[]
Family classes expose registered backend keys.
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
backends = VideoClipProcessor.list_available_backends()
print(backends)
["gstreamer", "ffmpeg", "moviepy"] # depends on registered backends
process(attachment, **kwargs)
async
Process a single attachment (or perform an aggregation on a list) and return the result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
T_in
|
The attachment or list of attachments to process. |
required |
**kwargs
|
Any
|
Additional keyword arguments forwarded to |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
T_out |
T_out
|
The result of the processing. |
Example
result = await processor.process(attachment)
print(result)
<processed attachment or transformed output>
process_batch(attachments, **kwargs)
async
Process a batch of attachments.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachments
|
list[T_in]
|
The batch of attachments to process. |
required |
**kwargs
|
Any
|
Additional keyword arguments forwarded to |
{}
|
Returns:
| Type | Description |
|---|---|
list[T_out]
|
list[T_out]: The result of the batch processing. |
Example
results = await processor.process_batch([attachment_1, attachment_2])
print(results)
[<result_1>, <result_2>]
ShotBasedSegmenter(detector=Detector.CONTENT, config=None)
Bases: BaseSegmenter[ShotBasedSegmenterConfig]
Segment video into shots with swappable decode + numpy/skimage scoring.
Frame decoding is delegated to the nested FrameDecodeProcessor family
(GStreamer by default; override via composite backend or
set_processor_backend); shot scoring runs on numpy + scikit-image.
Attributes:
| Name | Type | Description |
|---|---|---|
detector |
Detector
|
|
Initialize the FIPS-friendly content segmenter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Detector | str
|
Detector algorithm. Defaults to |
CONTENT
|
config
|
ShotBasedSegmenterConfig | BaseSegmenterConfig | dict[str, Any] | None
|
Scoring / decode configuration. Defaults to None (use defaults). |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
ValidationError
|
If scoring / decode knobs fail |
ImportError
|
If numpy / scikit-image / Pillow are not installed. |
config_model()
classmethod
Return this segmenter's concrete config model.
Returns:
| Type | Description |
|---|---|
type[ShotBasedSegmenterConfig]
|
type[ShotBasedSegmenterConfig]: The model used to validate shot-based segmenter configuration. |
ShotBasedSegmenterConfig
Bases: BaseSegmenterConfig
Configuration for ShotBasedSegmenter.
Decode and detector settings specific to content-based shot detection.
Attributes:
| Name | Type | Description |
|---|---|---|
sample_fps |
int | None
|
Decode sampling rate. Defaults to 5. |
target_width |
int | None
|
Decode downscale width. Defaults to 320. |
min_shot_duration |
float
|
Minimum accepted shot length in seconds. Defaults to 1.0. |
threshold |
float
|
Detector threshold. Defaults to 27.0. |
min_content_val |
float
|
|
window_width |
int
|
|
VideoClipProcessor()
Bases: BackendSelectableProcessor[Attachment, Attachment], ABC
Family base for clipping video attachments to time windows.
This class serves as a unified entry point for video clipping operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.
Why use this base class?
- Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
- Simplicity: No need to handle fallback logic or conditional imports yourself.
- Future-proofing: New backends can be added to the library without requiring changes to your application code.
Usage Example
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment
# Instantiates the best available backend automatically
processor = VideoClipProcessor.build()
# Set the target clipping window (start_time, end_time) in seconds
processor.set_windows([(10.0, 20.5)])
attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment
# Explicitly force the ffmpeg backend
processor = VideoClipProcessor.build(backend="ffmpeg")
processor.set_windows([(10.0, 20.5)])
attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment
# Explicitly force the moviepy backend
processor = VideoClipProcessor.build(backend="moviepy")
processor.set_windows([(10.0, 20.5)])
attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
set_windows(windows)
abstractmethod
Configure one or more [start, end] clipping windows for the next call.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
windows
|
list[tuple[float, float]]
|
List of (start, end) tuples in seconds. |
required |