Reactive Machines

LVSum: A Timestamp Benchmark for Long Video Summarization

Long video summarization presents significant challenges for large-scale linguistic models (MLLMs), especially in maintaining temporal fidelity over extended time periods and generating statistically and temporally based summaries. We present LVSum, a human-defined benchmark for evaluating long-form video summarization with fine-grained temporal alignment. LVSum has 72 different videos covering 13 domains with an average duration of 16 minutes, each with up to 10 human-generated tables containing temporal references. We conduct a comprehensive evaluation of the leading open source MLLMs using newly introduced LLM-based metrics for content and process compatibility, in addition to standard automated metrics. Our evaluation revealed three important findings: (1) transcripts contribute more to the quality of summaries than visual frames alone, (2) a significant performance gap persists between model-generated and human-written summaries, and (3) current MLLMs show systematic weaknesses in temporal support, adherence to guidelines, and cross-method compatibility.

Source link

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button