Skip to content

Overview

Media toolkit for processing video and audio attachments.

This package provides the infrastructure for media processing, including video clip extraction, frame sampling, audio extraction, keyframe extraction, and temporal segmentation.

Submodules

  • processor -- Processor families for video clip, frame sampling, audio extraction, deinterlace, frame extraction, and dense frame decode.
  • segmenter -- Temporal segmenters for splitting media into fixed-duration or shot-based chunks.
  • keyframe_extractor -- Keyframe extraction from video streams.

Usage

from gllm_multimodal.media_toolkit import processor, segmenter

AudioExtractionProcessor()

Bases: BackendSelectableProcessor[Attachment, Attachment], ABC

Family base for extracting audio tracks from video attachments.

This class serves as a unified entry point for audio extraction operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.

Why use this base class?

  • Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
  • Simplicity: No need to handle fallback logic or conditional imports yourself.
  • Future-proofing: New backends can be added to the library without requiring changes to your application code.

Usage Example

from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment

# Instantiates the best available backend automatically
processor = AudioExtractionProcessor.build()

attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment

# Explicitly force the ffmpeg backend
processor = AudioExtractionProcessor.build(backend="ffmpeg")

attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment

# Explicitly force the moviepy backend
processor = AudioExtractionProcessor.build(backend="moviepy")

attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)

BaseSegmenterConfig

Bases: BaseModel

Empty base config used to carry values into a concrete segmenter config.

Concrete segmenter configs declare and validate only the settings their implementation consumes. Extra fields are retained here so callers can up-cast an exact base config with from_dict.

from_dict(data=None) classmethod

Build a validated config from a dict, base config, or existing instance.

Sibling subclass configs are rejected: re-validating an unrelated model would silently drop (or carry over) fields the caller never set. Pass a dict or an exact BaseSegmenterConfig to up-cast into the concrete subclass.

Parameters:

Name Type Description Default
data Self | BaseSegmenterConfig | dict[str, Any] | None

Raw config payload. None uses field defaults. Defaults to None.

None

Returns:

Name Type Description
Self Self

Validated config of the concrete subclass.

Raises:

Type Description
TypeError

If data is a BaseModel that is neither an instance of cls nor an exact BaseSegmenterConfig, or is of any other unsupported type.

Detector

Bases: StrEnum

Supported shot-detector algorithm names.

FixedDurationSegmenter(config=None)

Bases: BaseSegmenter[FixedDurationSegmenterConfig]

Segment attachments using explicit per-segment durations.

segment returns cumulative time windows from config.segment_durations. materialize clips each window into a separate attachment using the cached VideoClipProcessor obtained from CompositeMediaMixin.

config_model() classmethod

Return this segmenter's concrete config model.

Returns:

Type Description
type[FixedDurationSegmenterConfig]

type[FixedDurationSegmenterConfig]: The model used to validate fixed-duration segmenter configuration.

materialize(attachment, segment, segment_index=0) async

Clip and return one attachment for a precomputed segment window.

This method is convenient when segment planning and clip extraction are performed in separate stages, and only selected windows should be materialized.

Parameters:

Name Type Description Default
attachment Attachment

Source media attachment to clip.

required
segment VideoSegment

Segment boundary plan to materialize.

required
segment_index int

Zero-based index used in generated output filenames. Defaults to 0.

0

Returns:

Name Type Description
Attachment Attachment

Clipped attachment with [VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment] metadata.

process(attachment, **kwargs) async

Materialize fixed-duration clips from one media attachment.

Unlike calling segment directly, this method returns real clipped attachment outputs with [VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment] metadata embedded on each result. It is the main runtime entrypoint when you need files/bytes for every configured duration window, not only boundary plans.

Parameters:

Name Type Description Default
attachment Attachment

Source media attachment to split.

required
**kwargs Any

Forwarded processing arguments accepted by the base media-toolkit contract.

{}
Notes
  1. Delegates shared validation and orchestration to MediaToolkit.process (inherited by BaseSegmenter).

Returns:

Type Description
list[Attachment]

list[Attachment]: One clipped attachment per configured segment window.

segment(attachment) async

Return computed fixed windows without creating clip attachments.

This is useful for previewing time boundaries (for inspection, logging, or downstream planning) before paying the cost of media clipping.

Parameters:

Name Type Description Default
attachment Attachment

Source media attachment. The payload itself is not read by this implementation when computing boundaries.

required
Notes
  1. Delegates shared validation to BaseSegmenter.segment.

Returns:

Type Description
list[VideoSegment]

list[VideoSegment]: Fixed cumulative windows derived from

list[VideoSegment]

segment_durations and start_time.

FixedDurationSegmenterConfig

Bases: BaseSegmenterConfig

Configuration for FixedDurationSegmenter.

Requires a non-empty segment_durations list.

Attributes:

Name Type Description
segment_durations list[float]

Ordered segment durations in seconds.

start_time float

Base start time for the first segment. Defaults to 0.0.

validate_segment_durations(value) classmethod

Reject non-positive segment durations.

Parameters:

Name Type Description Default
value list[float]

Candidate segment durations.

required

Returns:

Type Description
list[float]

list[float]: The durations unchanged when valid.

Raises:

Type Description
ValueError

If any duration is non-positive.

FrameDecodeProcessor()

Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC

Family base for dense video-frame decoding.

Why use this base class?

  • Portability: Swap FFmpeg vs GStreamer without changing call sites.
  • FIPS choice: backend="gstreamer" decodes with system plugins (no bundled-FFmpeg wheels); backend="ffmpeg" shells out to the system ffmpeg binary.
  • Separation: Frame decoding stays here; frame scoring (shot detection, keyframes) lives in segmenters/extractors.

Usage

from gllm_multimodal.media_toolkit.processor.frame_decode_processor import (
    FrameDecodeProcessor,
)

processor = FrameDecodeProcessor.build(backend="gstreamer")
frames = await processor.process(video_attachment)

frame_metadata(frame_index, fps)

Build the metadata dict attached to every decoded frame.

Parameters:

Name Type Description Default
frame_index int

Zero-based position in decode order.

required
fps float | None

Effective sampling rate, if known.

required

Returns:

Type Description
dict[str, Any]

dict[str, Any]: frame_index / timestamp / fps / frame_decode_backend mapping.

iter_rgb_frames_sync(attachment, *, sample_fps=None, target_width=None, as_rgb=True)

Yield decoded frames after the backend writes the full PNG sequence.

The decoder subprocess/pipeline completes first, so peak temp-disk usage is every sampled PNG at once. After that, this generator opens each file in order, yields the payload, and deletes the PNG so Python RAM stays O(1) in frames. Prefer this over process when callers only need a scored stream and can tolerate the peak-disk cost.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
sample_fps int | None

Override downsample rate. Defaults to the constructor config.

None
target_width int | None

Override downscale width. Defaults to the constructor config.

None
as_rgb bool

When True (default), yield RGB arrays. When False, yield raw PNG bytes (used by process).

True

Yields:

Type Description
tuple[Any, dict[str, Any]]

tuple[Any, dict[str, Any]]: (rgb_or_png_bytes, frame_metadata).

Raises:

Type Description
RuntimeError

If decoding fails or yields no frames.

FileNotFoundError

If a required decoder binary is missing.

FrameExtractionProcessor()

Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC

Family base for extracting image frames at timestamps.

Why use this base class?

  • Portability: Swap FFmpeg vs GStreamer without changing call sites.
  • Batch I/O: Extract many keyframes in one process call.
  • Separation: Keyframe planning stays in extractors; decode is here.

Usage

from gllm_multimodal.media_toolkit.processor.frame_extraction_processor import (
    FrameExtractionProcessConfig,
    FrameExtractionProcessor,
)

processor = FrameExtractionProcessor.build()
frames = await processor.process(
    video_attachment,
    process_config=FrameExtractionProcessConfig(timestamps=[1.5, 4.0]),
)

set_timestamps(timestamps) abstractmethod

Configure default timestamps used when process_config is omitted.

Parameters:

Name Type Description Default
timestamps list[float]

Non-empty list of non-negative seconds.

required

stream_frames(attachment, sample_fps, deinterlace=None) abstractmethod

Stream frames sampled at a fixed rate from a single decode pass.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
sample_fps float

Positive sampling rate in frames per second.

required
deinterlace DeinterlaceMode | None

Deinterlace override. Defaults to None.

None

Returns:

Type Description
AsyncIterator[tuple[float, ndarray]]

AsyncIterator[tuple[float, np.ndarray]]: Timestamps and BGR uint8 frames at source size.

FrameSamplingProcessor()

Bases: BackendSelectableProcessor[Attachment, Attachment], ABC

Family base for resampling video attachments to a target frame rate.

This class serves as a unified entry point for frame sampling operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.

Why use this base class?

  • Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
  • Simplicity: No need to handle fallback logic or conditional imports yourself.
  • Future-proofing: New backends can be added to the library without requiring changes to your application code.

Usage Example

from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import (
    FrameSamplingProcessor,
    GstFrameSamplingConfig,
)
from gllm_inference.schema import Attachment

# Instantiates the best available backend automatically
processor = FrameSamplingProcessor.build(
    config=GstFrameSamplingConfig(default_target_fps=2)
)

attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import FrameSamplingProcessor
from gllm_inference.schema import Attachment

# Explicitly force the ffmpeg backend
processor = FrameSamplingProcessor.build(backend="ffmpeg")

attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)

MediaToolkit()

Bases: ABC, Generic[T_in, T_out]

Base abstraction for all media toolkit processing components.

This class provides the shared lifecycle and registry behavior used by both: - concrete leaf processors (e.g. backend-specific audio/video processors), and - composite components (e.g. segmenters, keyframe extractors) that orchestrate nested processors.

Key responsibilities: - auto-register subclasses by class name for class-name-based construction via build; - provide consistent input validation against supported_mimetypes; - define async processing contracts through process and process_batch.

Contributor guidance: - inherit this class directly for concrete processors with custom behavior; - inherit BackendSelectableProcessor when one logical processor family maps to multiple backend implementations; - inherit composite bases (e.g. BaseSegmenter) for orchestration-style components.

Example
Building a processor by class name
from gllm_multimodal.media_toolkit.media_toolkit import MediaToolkit

processor = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
result = await processor.process(video_attachment)
Checking mimetype support
if processor.is_supported(attachment):
    result = await processor.process(attachment)
Listing registered processors
print(list(MediaToolkit.registry.keys()))
# ['GstAudioExtractionProcessor', 'GstVideoClipProcessor', ...]

Initialize processor logging.

name property

Return a stable component name for optional plan metadata.

Returns:

Name Type Description
str str

Component class name.

registry = {} class-attribute

Global class-name registry used by build.

Maps each registered subclass name to its concrete class type, enabling string-based construction such as MediaToolkit.build("AudioExtractionProcessor").

supported_mimetypes = ['*/*'] class-attribute

MIME types this processor accepts (supports wildcards, e.g. 'video/*').

Defaults to ['*/*'] (accept all). Override as a class attribute in subclasses.

__init_subclass__(**kwargs)

Register every concrete subclass into the global registry.

This hook is triggered automatically by Python whenever a class inherits from MediaToolkit (directly or indirectly). Registration happens at class definition/import time, so classes become immediately discoverable by build without manual setup.

Registration key
  1. The subclass' __name__ (e.g. "AudioExtractionProcessor").
  2. The value stored is the subclass type itself.
Why uniqueness is enforced
  1. build(class_name=...) uses this registry for class resolution.
  2. Duplicate class names would silently shadow earlier classes and could route builds to unintended implementations.
  3. To prevent that ambiguity, duplicate keys raise TypeError.

Parameters:

Name Type Description Default
**kwargs Any

Extra class declaration keyword arguments forwarded to parent __init_subclass__ implementations.

{}

Raises:

Type Description
TypeError

If a subclass with the same class name is already registered, preventing silent dispatch to the wrong implementation.

available_backends_for(class_name) classmethod

Return backend keys registered for a processor family class name.

Parameters:

Name Type Description Default
class_name str

Registered processor family class name.

required

Returns:

Type Description
list[str]

list[str]: Available backend keys. Empty when the class is unknown or not a backend-selectable family base.

Raises:

Type Description
ValueError

If the class name is unknown.

Example

backends = MediaToolkit.available_backends_for("VideoClipProcessor")
print(backends)
Output:
["gstreamer", "ffmpeg"]

build(class_name, backend=None, **kwargs) classmethod

Build a processor by class name.

Family abstract classes (e.g. AudioExtractionProcessor) resolve a concrete backend implementation via backend. Composite components (segmenters, keyframe extractors) store backend on the instance for nested resolution.

Parameters:

Name Type Description Default
class_name str

Registered subclass name.

required
backend str | MediaBackend | None

Backend key for family classes, or nested processor preference for composite instances. Defaults to None.

None
**kwargs Any

Constructor kwargs passed to the processor class.

{}

Returns:

Name Type Description
MediaToolkit MediaToolkit

Instantiated processor.

Raises:

Type Description
ValueError

If the class name is unknown.

Example

proc = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
seg = MediaToolkit.build(
    "FixedDurationSegmenter",
    backend="gstreamer",
    config={"segment_durations": [2.0]},
)
print(proc)
print(seg)
Output:
<GstAudioExtractionProcessor instance>
<FixedDurationSegmenter instance>

build_from_registry(backend=None, **kwargs) classmethod

Instantiate this registered class.

Subclasses override this hook to customize registry-based construction (e.g. backend-selectable families resolve a concrete backend; composites store backend for nested processor resolution).

Parameters:

Name Type Description Default
backend str | MediaBackend | None

Backend key forwarded to subclass overrides. Ignored by the base implementation.

None
**kwargs Any

Constructor kwargs passed to the processor class.

{}

Returns:

Name Type Description
MediaToolkit MediaToolkit

Instantiated processor.

Example

# Called indirectly by MediaToolkit.build(...)
processor = SomeRegisteredProcessor.build_from_registry(custom_flag=True)
print(processor)
Output:
<SomeRegisteredProcessor instance>

is_supported(attachment)

Return whether the attachment's mimetype is accepted by this processor.

Callers can use this to check compatibility before calling process or process_batch, avoiding a ValueError.

Parameters:

Name Type Description Default
attachment Attachment

The attachment to check.

required

Returns:

Name Type Description
bool bool

True if the attachment's mimetype matches any entry in supported_mimetypes (including wildcards). True is also returned when the attachment has no mimetype set.

Example

ok = processor.is_supported(Attachment(mime_type="video/mp4"))
print(ok)
Output:
True

list_available_backends() classmethod

Return backend keys when this class supports backend selection.

Returns:

Type Description
list[str]

list[str]: Available backend keys. Empty for classes that are not backend-selectable family bases.

Example

Base classes are not backend-selectable.

print(MediaToolkit.list_available_backends())
Output:
[]

Family classes expose registered backend keys.

from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor

backends = VideoClipProcessor.list_available_backends()
print(backends)
Output:
["gstreamer", "ffmpeg", "moviepy"]  # depends on registered backends

process(attachment, **kwargs) async

Process a single attachment (or perform an aggregation on a list) and return the result.

Parameters:

Name Type Description Default
attachment T_in

The attachment or list of attachments to process.

required
**kwargs Any

Additional keyword arguments forwarded to _process.

{}

Returns:

Name Type Description
T_out T_out

The result of the processing.

Example

result = await processor.process(attachment)
print(result)
Output:
<processed attachment or transformed output>

process_batch(attachments, **kwargs) async

Process a batch of attachments.

Parameters:

Name Type Description Default
attachments list[T_in]

The batch of attachments to process.

required
**kwargs Any

Additional keyword arguments forwarded to process.

{}

Returns:

Type Description
list[T_out]

list[T_out]: The result of the batch processing.

Example

results = await processor.process_batch([attachment_1, attachment_2])
print(results)
Output:
[<result_1>, <result_2>]

ShotBasedSegmenter(detector=Detector.CONTENT, config=None)

Bases: BaseSegmenter[ShotBasedSegmenterConfig]

Segment video into shots with swappable decode + numpy/skimage scoring.

Frame decoding is delegated to the nested FrameDecodeProcessor family (GStreamer by default; override via composite backend or set_processor_backend); shot scoring runs on numpy + scikit-image.

Attributes:

Name Type Description
detector Detector

content | adaptive | threshold.

Initialize the FIPS-friendly content segmenter.

Parameters:

Name Type Description Default
detector Detector | str

Detector algorithm. Defaults to content.

CONTENT
config ShotBasedSegmenterConfig | BaseSegmenterConfig | dict[str, Any] | None

Scoring / decode configuration. Defaults to None (use defaults).

None

Raises:

Type Description
ValueError

If detector is not a supported algorithm name.

ValidationError

If scoring / decode knobs fail ShotBasedSegmenterConfig.

ImportError

If numpy / scikit-image / Pillow are not installed.

config_model() classmethod

Return this segmenter's concrete config model.

Returns:

Type Description
type[ShotBasedSegmenterConfig]

type[ShotBasedSegmenterConfig]: The model used to validate shot-based segmenter configuration.

ShotBasedSegmenterConfig

Bases: BaseSegmenterConfig

Configuration for ShotBasedSegmenter.

Decode and detector settings specific to content-based shot detection.

Attributes:

Name Type Description
sample_fps int | None

Decode sampling rate. Defaults to 5.

target_width int | None

Decode downscale width. Defaults to 320.

min_shot_duration float

Minimum accepted shot length in seconds. Defaults to 1.0.

threshold float

Detector threshold. Defaults to 27.0.

min_content_val float

adaptive-only absolute score floor. Defaults to 15.0.

window_width int

adaptive-only neighbour half-window. Defaults to 2.

VideoClipProcessor()

Bases: BackendSelectableProcessor[Attachment, Attachment], ABC

Family base for clipping video attachments to time windows.

This class serves as a unified entry point for video clipping operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.

Why use this base class?

  • Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
  • Simplicity: No need to handle fallback logic or conditional imports yourself.
  • Future-proofing: New backends can be added to the library without requiring changes to your application code.

Usage Example

from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment

# Instantiates the best available backend automatically
processor = VideoClipProcessor.build()

# Set the target clipping window (start_time, end_time) in seconds
processor.set_windows([(10.0, 20.5)])

attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment

# Explicitly force the ffmpeg backend
processor = VideoClipProcessor.build(backend="ffmpeg")
processor.set_windows([(10.0, 20.5)])

attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment

# Explicitly force the moviepy backend
processor = VideoClipProcessor.build(backend="moviepy")
processor.set_windows([(10.0, 20.5)])

attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)

set_windows(windows) abstractmethod

Configure one or more [start, end] clipping windows for the next call.

Parameters:

Name Type Description Default
windows list[tuple[float, float]]

List of (start, end) tuples in seconds.

required