Processing hundreds of hours of conference talks manually is one of the most resource-intensive operational challenges for event platforms like FOSSASIA Eventyay. During Google Summer of Code 2026, we set out to solve this problem by architecting VEditor—a standalone, high-throughput video review and transcoding pipeline.

[Raw Ingest] -> [PyAV Demuxer] -> [Silence/Slide Detect] -> [Human Review Gate] -> [RQ Transcoder] -> [Final 1080p/AV1]

Why Subprocess Shelling Fails at Scale

The standard naive approach to video automation in Python involves executing ffmpeg via subprocess.Popen:

# The Subprocess Pitfall
subprocess.run([
    "ffmpeg", "-i", input_path, "-ss", start_time, "-to", end_time, "-c", "copy", output_path
])

While simple for one-off conversions, subprocess shelling breaks down when executing fine-grained analysis:

  1. Process Fork Overhead: Spawning hundreds of sub-shells creates heavy OS process thrashing.
  2. Brittle String Parsing: Extracting PTS timestamps, keyframe indexes, and silence intervals from stdout logs requires error-prone regex parsing.
  3. No Direct Frame Memory Access: Analyzing luminance drops or audio waveforms requires piping raw YUV/PCM streams over stdin/stdout pipes.

The Solution: Native C-Bindings with PyAV

By utilizing PyAV, Python interfaces directly with FFmpeg’s underlying libavformat, libavcodec, and libavutil C libraries:

import av

def detect_talk_boundaries(video_path: str, silence_threshold_db: float = -35.0):
    container = av.open(video_path)
    audio_stream = next(s for s in container.streams if s.type == 'audio')
    
    cue_points = []
    for frame in container.decode(audio_stream):
        # Directly inspect audio frame amplitude array in memory
        nd_array = frame.to_ndarray()
        rms = (nd_array ** 2).mean() ** 0.5
        if rms > silence_threshold_db:
            cue_points.append(frame.pts * audio_stream.time_base)
            
    container.close()
    return cue_points

This yields zero intermediate disk I/O, sub-millisecond timestamp inspection, and memory-safe frame navigation.

Distributed Orchestration with Redis Queue (RQ)

Heavy transcoding jobs must never block user-facing APIs. We decomposed VEditor into:

  • FastAPI Gateway: Handles upload presigned URLs, session metadata, and review gate decisions.
  • Worker Pools: Redis Queue (RQ) workers listening on separate queues (analysis, preview, transcode) with isolated worker concurrency limits based on available CPU cores.
from rq import Queue
from redis import Redis

redis_conn = Redis(host="redis", port=6379)
analysis_q = Queue("analysis", connection=redis_conn)
transcode_q = Queue("transcode", connection=redis_conn)

@app.post("/talks/{talk_id}/ingest")
async def ingest_talk(talk_id: str, payload: IngestRequest):
    # Enqueue analysis job non-blockingly
    job = analysis_q.enqueue("workers.tasks.analyze_stream", talk_id, payload.s3_uri)
    return {"status": "enqueued", "job_id": job.id}

The Human-in-the-Loop Review Gate

Fully automated trimming inevitably encounters edge cases (e.g., audio microphone checks before a talk begins). Rather than blindly publishing auto-trimmed videos, the pipeline introduces a Review State Gate:

UPLOADED -> ANALYZING -> PENDING_REVIEW -> [Staff Approve/Adjust] -> TRANSCODING -> PUBLISHED

Reviewers are presented with a lightweight timeline preview showing auto-detected speech onsets, slide transition markers, and low-res proxy clips. Once approved, the high-bitrate AV1/H.264 transcode task is enqueued.

By combining low-level C bindings, asynchronous task isolation, and pragmatic human oversight, VEditor reduces conference post-production turnaround from weeks to hours.