Laptop on a desk displaying the captioned video player with subtitle language options for a panel discussion.

The Modern Livestream is More Than Just Video

For years, a livestream player had one basic job: play the video.

It was simple: open the link, click play, watch the event. That was the whole experience people expected, with nothing more and only one language.

However, when accessibility and multilingual communication are built into the player, the same live event can provide an audience with so much more. During a broadcast, a modern player can:

Same event, same link, but four different ways to experience it, all chosen by the viewer.

So, what makes a streaming player accessible? What should your organization look for when choosing one?

Well, that is what this guide is all about!

Does an Accessible Player Matter?

Picture your event: a commencement, conference, service, or town hall. This event will reach many different people online. Do you know who is on the other end watching? What are their needs? What would make their experience easier?

Unless a survey went out ahead of time, this can be very hard to determine, but you can expect that at least one viewer will need one or more accessibility services. With an accessible player, each viewer can choose how they take in your content:

The point isn’t to include a single feature. It’s to offer multiple options and allow the viewer to decide. This ensures that no one in the audience is left out.

That is what accessibility looks like when it is built in from the start rather than added as an afterthought.

Aberdeen Technical University Commencement livestream demo

Try it for yourself. Use the controls in the player to switch between captions, subtitle languages, and translated audio. Each viewer can make their own selections while watching the same live event.

“Live” Stream – Not VOD

It is worth being clear about the current challenge. Many conventional streaming platforms already have accessibility and language options for recorded or on-demand (VOD) video. Viewers can experience recorded video with captions and translated audio tracks on these platforms.

Livestreaming is a different story. During a live event, many of the capabilities offered for VOD are more limited. This is the gap we aimed to bridge with an accessibility-focused player built to bring these options into the live viewing experience.

Here is a look at how a purpose-built player compares to what mainstream platforms offer during a live event:

Capability
Advantage Accessibility-Focused Live Player
Mainstream Mainstream Live Player
Live Captions
Accessibility-Focused Yes – human or automated, embedded during the event
Mainstream Yes (generally) – modern stacks include automated captions that vary in accuracy
Translated Subtitles
Accessibility-Focused Yes – viewers can pick their own language live
Mainstream No (typically) – if offered, the choices are very limited
Translated Audio Tracks (Dubbing)
Accessibility-Focused Yes – viewers can pick their own language live
Mainstream No – very rare if offered during a live event and accuracy varies
Picture-in-Picture ASL
Accessibility-Focused Yes – a signing window is synced with the main video
Mainstream No – if offered, the producer has this burned into the content
Viewer Choice
Accessibility-Focused Yes – each viewer can select their preferred accessibility option
Mainstream Limited – generally the same feed is served to everyone
Live Synchronization
Accessibility-Focused Yes – video, captions, translations, and audio are all aligned
Mainstream Varies – live synchronization across multiple tracks can be difficult

Reliability & Responsibility

None of these services matter if the stream itself is not dependable. Viewers need video that plays without constant buffering, captions and translated subtitles that arrive on time, and translated audio that is clear and synchronized with the event.

It is the responsibility of the streaming provider to have a reliable foundation that can continually play out the stream. On top of this foundation is a capable live player built to deliver:

Behind the scenes, the player adjusts to each viewer’s connection in real time and keeps each track locked to the same clock. Keeping video, captions, translations, and multiple audio tracks synchronized is one of the most important parts of delivering a seamless live experience. A good player handles that synchronization so the viewer never has to think about it.

Advanced Technology ≠ Complicated Setup

If this all sounds complicated, don’t worry. The complexity of the technology stays behind the pipeline and player. At Aberdeen, we walk clients through the whole process:

Your team can stay focused on the event while we handle the work behind the scenes that makes the stream accessible.

One Live Event, Multiple Ways to Participate

The same livestream can be experienced in different ways. Each viewer selects the accessibility and language options that work best for them.

Diagram showing one live event branching into English captions, Spanish subtitles, Mandarin audio, and ASL interpretation.

What Should You Look for in an Accessible Streaming Player?

Here is a practical checklist if you are evaluating options. These questions are worth asking any provider and touch on everything above:

Engineering Should Be Invisible to the Viewer

With all the engineering underneath, the viewer should never have to think about it. No one should need a technical background in adaptive bitrate streaming or audio track management to experience a graduation or service. They should simply open the link, choose the language or accessibility option that works for them, and enjoy the event.

That is the real goal: technology sophisticated enough to include everyone and simple enough that no one notices it is there.

Continue the Conversation

Make Your Next Livestream More Accessible

Every audience is different. Giving viewers access to captions, translated subtitles, live language audio, and ASL within the same viewing experience lets more people participate in the way that works best for them. Talk with Aberdeen about building the right accessible livestream experience for your next event.
Exterior of the U.S. Department of Justice building with columns and the “Department of Justice” inscription across the façade

For many organizations, the April 24, 2026, deadline around ADA Title II has raised an important question: Is this a new requirement, or something that’s been in place all along?

The answer is straightforward: Accessibility under ADA Title II is not new. What’s new is clarity.

In 2024, the Department of Justice issued a final rule that formally defines how ADA Title II applies to websites, mobile apps, and digital content. For the first time, public entities now have a clear technical standard and a firm deadline.

This post breaks down:

What ADA Title II Already Requires

Under the Americans with Disabilities Act (ADA) Title II, state and local governments have long been required to provide equal access to their programs, services, and activities, along with effective communication for individuals with disabilities.

In practice, this has always applied to core public functions like meetings, educational programs, and government services. As digital communication became central to how these services are delivered, enforcement made it increasingly clear that the same expectations extended to websites, online video, and other digital materials.

Accessibility in digital environments wasn’t new; it was a continuation of an existing requirement.

What the 2024 Rule Clarifies

The DOJ’s 2024 update does not change the core obligation. It defines it.

For the first time, public entities now have:

Why Clarification Was Needed

While the requirement itself was well established, how to meet it was not.

There was no officially defined technical standard, no universal deadline, and no consistent enforcement model. As a result, organizations relied on interpretation, guidance, and precedent to determine what “accessible” meant in practice.

Much of that guidance came through enforcement and legal action. The Department of Justice and the Office for Civil Rights investigated complaints and entered into resolution agreements, while high-profile cases helped shape expectations. The University of California, Berkeley case required the removal or remediation of inaccessible online video content, and lawsuits involving Harvard and MIT reinforced expectations around captioning and digital access.

These cases made one thing clear: Accessibility was required, but organizations didn’t have a consistent, measurable way to implement it.

ADA Title II: Before vs. Now

Category
Previously Before 2024 Rule
Now 2024 Rule – Effective 2026/2027
Legal Requirement
Previously Accessibility required under ADA Title II
Now Accessibility still required
Digital Coverage
Previously Implied through interpretation and case law
Now Explicitly includes websites, apps, and digital content
Technical Standard
Previously Not formally defined
Now WCAG 2.1 Level AA required
Enforcement Style
Previously Complaint-driven (OCR, lawsuits)
Now Proactive and enforceable
Deadlines
Previously No universal deadline
Now April 2026 / April 2027
Captioning Expectation
Previously Required under "effective communication"
Now Clearly required under WCAG
Consistency
Previously Varied by organization
Now Standardized across public entities

What the Rule Says About Captions

One of the most immediate impacts of the rule is clarity around captions.

Under WCAG 2.1 Level AA:

This aligns with how accessibility has already been enforced, but now it is explicitly defined and expected. Just as important, the standard is not simply whether captions exist, it’s whether they are effective.

WCAG does not define a specific accuracy percentage. Instead, it requires that captions present the full meaning of the content, including spoken dialogue and relevant non-speech elements, in a way that is properly synchronized and easy to follow. This is reinforced by ADA Title II’s broader requirement for effective communication: Captions must allow a viewer to fully understand the message—not just approximate it.

In practice, that means:

With that in mind, it’s important to understand how different captioning approaches align with these expectations. There are two primary approaches used today: automated captioning powered by AI (ASR) and human captioning performed by trained writers.

In practice, the right choice comes down to context. ASR can be effective in controlled environments with clear audio and lower risk, offering a scalable and cost-efficient solution. Human captioning is better suited for high-stakes, complex, or public-facing content where accuracy, speaker identification, and reliability are critical.

The goal isn’t choosing a method—it’s ensuring the message is fully understood.

The “Older Content” Exception

The rule includes a limited exception for content created before April 24, 2026, but it’s narrower than many expect.

Older content can remain as-is only if it is truly archival. That means it is not actively used, not updated, and not part of any current program, service, or activity.

Where this gets important is how “use” is defined. If older content is still being used in any meaningful way, it must be made accessible—even if it was created years ago.

Common scenarios where older content must be updated

When content may remain exempt

Content may qualify for the exception if it is:

The practical way to think about it: If your audience is expected to use it, it must be accessible. The exception isn’t based on age; it’s based on relevance and use.

What About Language Requirements?

The rule does not require translation. There is no percentage threshold that triggers multilingual content, no requirement to offer multiple languages, and no rule that content in one language must be mirrored in another.

What the rule does require is consistency: Any language you provide must be accessible.

If an organization offers content in Spanish, that version must be accessible in Spanish. If content is delivered in English, it must be accessible in English.

For example, a Spanish video would need Spanish captions, and an English livestream would need real-time English captions.

Language access itself is governed by other regulations. ADA Title II focuses specifically on accessibility for people with disabilities, ensuring that whatever content is provided can be fully understood.

Who Needs to Comply, and When?

The rule applies to public entities under ADA Title II, including state and local governments, public universities, school systems, and municipal agencies.

The timeline is based on population size:

April 24, 2026

Applies to public entities serving 50,000+ people, such as:

April 26, 2027

Applies to entities serving under 50,000 people, including:

The requirement is the same for both groups—only the timeline differs.

What About Churches and Private Organizations?

Churches are not considered public entities under ADA Title II and are generally exempt from ADA Title III as well. This means they are not legally required to meet WCAG standards.

However, many churches are still adopting accessibility tools like captions and translation—not because they are required to, but because they recognize the value. Accessibility improves understanding, supports multilingual communities, and helps remove barriers for first-time visitors.

Accessibility in this context isn’t about compliance. It’s about connection.

Final Takeaway

ADA Title II has required accessibility for decades. The 2024 rule does not introduce a new obligation—it provides a clear, consistent framework for meeting one that already existed.

For public entities, that means:

Accessibility has always been about ensuring people can fully receive the message. Now, there is a clear path for how to deliver it.

Closed captions have long been one of the most important accessibility tools in modern media. They transform spoken words, music, and sound cues into text, allowing Deaf and hard-of-hearing audiences to fully engage with film, television, and live programming.

Yet despite decades of progress in video technology, caption presentation itself has changed very little. The familiar format of white text at the bottom of the screen remains the standard across most platforms. While effective, this approach often strips away elements that hearing viewers naturally perceive, such as tone, pacing, emphasis, and speaker identity.

That is why we were intrigued when we came across Caption with Intention.

Moving Beyond Static Text

Caption with Intention is not simply a visual redesign. It is a captioning design system built on a simple yet powerful premise: captions should convey not only what is said but also how it is said.

The project explores ways to represent aspects of speech that traditional captions rarely capture:

These ideas may sound subtle, but they address a real gap in how captioned content is experienced.

Designed With Community Input

One of the most compelling aspects of Caption with Intention is its development process. The system was shaped through collaboration with members of the Deaf and hard-of-hearing community. That involvement grounds the work in lived experience rather than purely aesthetic experimentation.

Accessibility innovations are most meaningful when they are informed by the audiences they are intended to serve. This project reflects that principle.

A Glimpse of a Possible Future

Caption with Intention is currently a design framework rather than a fully automated captioning engine. Its concepts still require thoughtful implementation. Even so, the direction is notable.

As AI, real-time rendering, and player technologies continue to evolve, it is easy to imagine a future where expressive captioning systems like this can be applied at scale. Not as decorative features, but as standard components of accessible storytelling.

Why It Caught Our Attention

At Aberdeen, we spend a great deal of time thinking about how captions function in practical, regulatory, and production contexts. Accuracy, timing, compliance, and reliability always come first.

Projects like Caption with Intention invite a different but equally important question:
What if captions could better reflect the emotional and narrative texture of a scene?

The idea does not replace the fundamentals. It expands the conversation.

We are not involved in the project, nor are we presenting this as an endorsement or partnership. We simply find the thinking behind it compelling. It represents the kind of experimentation that can influence how accessibility evolves over time.

We will be watching its progress with interest and testing its ideas when the technology and workflows feel ready.

Because the future of captions is not only about access. It is also about experience.

Gather25 live broadcast interface showing a female speaker on stage, with multilingual audio options in 84 languages including Spanish, Russian, and Swahili.

Gather25, hosted by IF:Gathering, was a groundbreaking 25-hour global broadcast connecting audiences across every continent. With over 1.25 million online viewers and participation from more than 21,000 Gather Groups worldwide, the event aimed to inspire and unite the global church community.

Aberdeen Broadcast Services has partnered with IF:Gathering since 2020, providing both live captioning for streaming events and post-produced captioning for archived content. Gather25 marked an ambitious new chapter in that partnership—the first time Aberdeen was brought on to help scale the event’s global reach through real-time translation.

As far as we know, this was a first-of-its-kind undertaking: a multilingual livestream of this magnitude, requiring impeccable coordination across dozens of languages and platforms. Aberdeen was one of several trusted vendors, working alongside the technology and broadcast partners hired for the event to make it all possible.

The Challenge: Delivering 25 Hours of Multilingual, Real-Time Access

Supporting a continuous, multilingual broadcast of this scale introduced several technical challenges:

The Solution: Continuous Live Captioning and Translation Across 84 Languages

Before the event went live, a workflow that could support uninterrupted, real-time captioning and translation across dozens of simultaneous streams needed to be built. This required deep coordination with multiple teams to align audio sources, language feeds, and delivery endpoints. Our goal was to ensure every segment of the broadcast could be accurately captioned and translated with minimal manual intervention once the event began.

A key technology partner behind this workflow was SyncWords, whose platform Aberdeen leverages to manage real-time captioning and translation delivery. SyncWords played a vital behind-the-scenes role in not only powering the infrastructure we used to deploy captions across dozens of streams but also collaborating directly with engineers at Sardius and Elemental Media to implement specialized audio-isolation coding for the event. This coordination ensured that every audio feed we received was optimized for clean, accurate transcription and translation at scale.

Pre-Event Testing

To guarantee performance, Aberdeen conducted over 50 hours of pre-event testing, stress-testing multi-language streams, and simulating 25-hour sessions to ensure system endurance. A key priority was ensuring that VTT caption files would function as continuously updated feeds, rather than static uploads. This real-time updating was essential for Sardius’ platform to support both live captions and rolling DVR features, while also allowing seamless access to captions during on-demand playback after the event.

Streaming Caption Workflow

In a typical video workflow, captions are created after recording, uploaded separately, and synced to on-demand content. For Gather25’s livestream, captions had to be generated and delivered in real time, alongside the video stream.

The system worked like this:

This architecture required precise timing, structured file delivery, and full alignment with Sardius’ streaming infrastructure.

What's an HLS Stream?

An HLS stream (HTTP Live Streaming) delivers video content over the internet in small, manageable chunks. The video is split into short segments (usually 2–10 seconds long) and saved as .ts (transport stream) files. A playlist file (called a .m3u8) tells the video player what order to play those chunks in. As a viewer watches, their device downloads and plays the segments one at a time, allowing smooth playback, even with slow or fluctuating internet.

Workflow Optimization

Aberdeen’s team engineered a sophisticated workflow to process 20 incoming HLS feeds. Working closely with Element Media Group, which managed master control and delivered stripped audio to Sardius, Aberdeen received feeds prepped for accessibility:

Diagram of Gather25 broadcast workflow with human captions, ASR translation, and 84 multilingual streams for TV and live streaming.

Impact: Real-Time Engagement and Global Reach Through Accessibility

Gather25’s accessibility efforts produced measurable results:

By combining AI-driven automation, human captioning expertise, and a deep integration with broadcast systems, Aberdeen Broadcast Services delivered scalable, high-quality accessibility at a truly global level.

Conclusion: Advancing Live Event Accessibility on a Global Scale

Gather25’s mission to unite believers around the world was made stronger through its commitment to accessibility. With real-time captioning and translation across 84 language streams, Aberdeen Broadcast Services helped make this global event inclusive, impactful, and available to all.

For organizations planning large-scale, multilingual broadcasts, Aberdeen’s tested and proven solutions offer the reliability and scalability needed to reach a worldwide audience.

Let’s talk about how we can support your next event. Contact us to learn more.

Closed captioning serves as a powerful tool that extends its impact far beyond aiding the deaf and hard-of-hearing community. Its significance transcends age, abilities, and background, making it an invaluable resource for both educators and learners. In the digital age, closed captioning has emerged as a transformative resource, with research revealing that students, English language learners, and children with learning disabilities who watch programs with closed captioning turned on improve their reading skills, increase their vocabulary, and enhance their focus and attention.

The scholarly article, Closed Captioning Matters: Examining the Value of Closed Captions for All Students (Smith 231) states that “Previous research shows that closed captioning can benefit many kinds of learners. In addition to students with hearing impairments, captions stand to benefit visual learners, non-native English learners, and students who happen to be in loud or otherwise distracting environments. In remedial reading classes, closed captioning improved students’ vocabulary, reading comprehension, word analysis skills, and motivation to learn (Goldman & Goldman, 1988). The performance of foreign language learners increased when captioning was provided (Winke, Gass, & Sydorenko, 2010). Following exams, these learners indicated that captions lead to increased attention, improved language processing, the reinforcement of previous knowledge, and deeper understanding of the language. For low-performing students in science classrooms, technology-enhanced videos with closed captioning contributed to post-treatment scores that were similar to higher-performing students (Marino, Coyne, & Dunn, 2010). The current findings support previous research and highlight the suitability of closed-captioned content for students with and without disabilities.”

Reading Rockets, a national public media literacy initiative provides resources and information on how young children learn and how educators can improve their students’ reading abilities. In the article, Captioning to Support Literacy, Alise Brann confirms that “Captions can provide struggling readers with additional print exposure, improving foundational reading skills.”

She states, “In a typical classroom, a teacher may find many students who are struggling readers, whether they are beginning readers, students with language-based learning disabilities, or English Language Learners (ELLs). One motivating, engaging, and inexpensive way to help build the foundational reading skills of students is through the use of closed-captioned and subtitled television shows and movies. These can help boost foundational reading skills, such as phonics, word recognition, and fluency, for a number of students.”

Research clearly demonstrates that “people learn better and comprehend more when words and pictures are presented together. The combination of aural and visual input gives viewers the opportunity to comprehend information through different channels and make connections between them” (The Effects of Captions on EFL Learners’ Comprehension of English-Language Television Programs).

From bolstering reading skills, to enhancing focus and language comprehension, the benefits of closed captioning are numerous. We at Aberdeen Broadcast Services are committed to providing quality closed captions for television (TV) and educational programming.

Here is the public service announcement (PSA) we released in 2016 on local broadcast stations, emphasizing how closed captioning can enhance children's literacy skills.

Photo of a hand on a remote scrolling through a video library

Been tasked with figuring out how to implement closed captions in your video library? The process can be overwhelming at first. While evaluating closed captioning vendors, it’s good to understand the benefits of captioning, who your audience is, what to consider when it comes to quality, and what to expect from a vendor.

There are several things that an organization should consider and evaluate before choosing a closed captioning vendor. Some of the most important factors include:

Benefits of Closed Captioning

Overall, closed captioning is a valuable tool that can benefit a wide range of audiences. It makes videos more accessible, engaging, and comprehensible for everyone.

Evaluating Vendors

By considering these factors, organizations can choose a closed captioning vendor that will meet their needs and provide a high-quality service:

What to Expect in the Process

Use these tips when evaluating closed captioning vendors and you’ll ensure that their videos are accessible to everyone and that they provide a positive viewing experience for all viewers.

Photo of a conference call on Zoom

On October 11, 2022, the Federal Communications Commission (FCC) released the latest CVAA biennial report to Congress, evaluating the current industry compliance as it pertains to Sections 255, 716, and 718 of the Communications Act of 1934. The biennial report is required by the 21st Century Communications and Video Accessibility Act (CVAA), which amended the Communications Act of 1934 to include updated requirements for ensuring the accessibility of "modern" telecommunications to people with disabilities.

FCC rules under Section 255 of the Communications Act require telecommunications equipment manufacturers and service providers to make their products and services accessible to people with disabilities. If such access is not readily achievable, manufacturers and service providers must make their devices and services compatible with third-party applications, peripheral devices, software, hardware, or consumer premises equipment commonly used by people with disabilities.

Accessibility Barriers

Despite major design improvements over the past two years, the report reveals that accessibility gaps still persist and that industry commenters are most concerned about equal access on video conferencing platforms. The COVID-19 pandemic has highlighted the importance of accessible video conferencing services for people with disabilities.

Zoom, BlueJeans, FaceTime, and Microsoft Teams have introduced a variety of accessibility feature enhancements, including screenreader support, customizable chat features, multi-pinning features, and “spotlighting” so that all participants know who is speaking. However, commentators have expressed concern over screen share and chat feature compatibility with screenreaders along with the platforms’ synchronous automatic captioning features.

Although many video conferencing platforms now offer meeting organizers synchronous automatic captioning to accommodate deaf and hard-of-hearing participants, the Deaf and Hard of Hearing Consumer Advocacy (DHH CAO) pointed out that automated captioning sometimes produces incomplete or delayed transcriptions and even if slight delays of live captions cannot be avoided, these captioning delays may cause “cognitive overload.” Comprehension can be further hindered if a person who is deaf or hard of hearing cannot see the faces of speaking participants, for “people with hearing loss rely more on nonverbal information than their peers, and if a person misses a visual cue, they may fall behind in the conversation.”

Automated vs. Human-generated Captions

At present, the automated captioning features on these conference platforms have an error rate of 5-10%. That’s 5-10 errors per 100 words spoken and when the average conversation rate of an English speaker is 150 words per minute, you’re looking at the possibility of over a dozen errors a minute.

Earlier this year, our team put Adobe’s artificial intelligence (AI) powered speech-to-text engine to the test. We tasked our most experienced Caption Editor with using Adobe’s auto-generated transcript to create & edit the captions to meet the quality standards of the FCC and the deaf and hard of hearing community on two types of video clips: a single-speaker program and one with multiple speakers.

How did it go? Take a look: Human-generated Captions vs. Adobe Speech-to-text

Open captions and closed captions are both used to provide text-based representations of spoken dialogue or audio content in videos, but they differ in their visibility and accessibility options.

Here's the difference between closed and open captions:

Open Captions

Closed Captions

FeatureOpen CaptionClosed Captions
VisibilityPermanently embedded in the videoSeparate text track that can be turned on or off
AccessibilityCannot be turned offCan be turned on or off by the viewer
ApplicationsWide audiences, noisy environmentsDiverse audiences, compliance with accessibility regulations
CreationAdded during video productionGenerated in real-time or embedded manually during post-production or uploaded as a sidecar file

Both open and closed captions serve the purpose of making videos accessible to individuals who are deaf or hard of hearing, those who are learning a new language, or those who prefer to read the text alongside the audio.

The choice between open or closed captions depends on the specific requirements and preferences of the content creators and the target audience.

In the July ‘21 release of Premiere Pro, Adobe introduced its artificial intelligence (AI) powered speech-to-text engine to help creators make their content more accessible to their audiences. Their extensive toolset allows their users to edit, stylize, and export captions in all supported formats straight out of the sequence timeline of a Premiere Pro project. A 3-step process of auto-transcribing, generating, and stylizing captions all within the platform already familiar to its users delivers a seamless experience from beginning to end. But how accurate is the final product?

This blog article was published in March 2022, and since then, ASR (Automatic Speech Recognition) technology has advanced significantly. While AI-powered ASR still does not outperform human writers — which we firmly consider the gold standard — these advancements have been so substantial and continue to improve. This progress gives us the confidence to use ASR as a budget-friendly alternative for specific applications.

Learn more about our ASR services here:

Today, at their best, AI captions have an error rate of 5-10% - much improved over the 80% accuracy we saw just a few years ago. High accuracy is crucial for the deaf and hard-of-hearing audience as each error adds to the possibility of confusing the message. To protect all audiences that rely on captioning to understand television programming, the Federal Communications Commission (FCC) set a detailed list of quality standards that all captions must meet to be acceptable for broadcast back in 2015. Preceding those standards, the Described and Captioned Media Program (DCMP) published its Captioning Key manual over 20 years ago and has since been a valuable reference for captioning of both entertainment and educational media targeted to audiences of all age groups. Simply having captions present on your content isn’t enough, it needs to be accurate and best replicate the experience for all audiences.

Adobe’s speech-to-text engine has been one of the most impressive that our team has seen to date, so we decided to take a deeper look at it and run some tests. We tasked our most experienced Caption Editor with using Adobe’s auto-generated transcript to create & edit the captions to meet the quality standards of the FCC and the deaf and hard of hearing community on two types of video clips: a single-speaker program and one with multiple speakers. Our editor used our Pop-on Plus+ caption product for these examples, which are our middle-tier quality captions that fulfill all quality standard requirements but are not always 100% free of errors.

Did using Adobe’s speech-to-text save time, or did it create more work in the editing process than needed? Here’s how it went…

In-depth comparison documents that evaluate the captions cell-by-cell are available for download here:

Single Speaker Clip

In this example, we used the perfect scenario for AI: clear audio, a single speaker at an optimal words-per-minute (WPM) speaking rate, and no sound effects or music.

The captions contained the following issues that would need to be corrected by the Caption Editor:

Here’s the clip with Adobe’s speech-to-text captions overlayed on the top half of the video, and ours on the bottom half.

Multiple Speaker Clip

For the next clip, we went with a more realistic example of television programming where there are multiple speakers, an area where AI is known to struggle and has difficulties identifying the speakers. This clip also features someone with a pronounced accent, commentators speaking over one another, and proper names of athletes – all of which our editors take the time to research and understand.

The same errors detailed in the single-speaker example are present throughout, among the other difficulties we expected it to have. In fact, there were so many errors that our editor was unable to use the transcript from Adobe and started from the beginning using our own workflow.

Here’s a sample of the first 9 cells of captions with what Adobe transcribes in the first column, notes from our Caption Editor, and how it should look.

Adobe’s Automated SRT Caption FileIssueFormatted by Aberdeen
something
 you are never seen in your life, correct?
No speaker ID.(Pedro Martinez)
It's something you have
never seen in your life,
“Correct” is spoken by new speaker.(Matt Vasgersian)
Correct!
So it's.Missing text.So it's--so it's MVP
of the year!
So we're all watching something
 different. OK
(Pedro)
We're all watching
something different.
He gets the MVP.Okay, he gets the MVP.
I'd be better off.Completely misunderstood music lyrics.♪ Happy birthday to you ♪
Oh, you, you guys.(Matt)
You guys.
Let me up here to dove into the opening
night against the Hall of Fame.
Merged multiple sentences together.Just left me up here to die.
You left me up here to die
against the hall of famer.

Take a look at the clip. Again, with Adobe's speech-to-text on the top and Aberdeen on the bottom.

In-depth comparison documents that evaluate the captions cell-by-cell are available for download here:

The Verdict

Overall, the quality of the auto-generated captions exceeded expectations, and we found them to be in the top tier of speech-recognition engines available. The timing and punctuation were particularly impressive. However, when doing a true comparison to the captioning work that we would consider acceptable, AI does not meet Aberdeen’s broadcast quality standard.

Aberdeen's post-production Caption Editors are detail-oriented and grammar-savvy and always strive to portray every element of the program with 100% accuracy so that the viewer misses nothing. For our most experienced Caption Editor, it took a 5:1 ratio in time for them to edit and correct the single-speaker clip; meaning, for every minute of video, it took 5 minutes to clean up the transcript and captions. Assuming your team is educated in the proper timing of caption cells, line breaks, and grammar, a 30-minute program may take over 2.5 hours to bring up to standards with a usable transcript. In the second example, the transcript was unusable and would have taken more time to clean up than it did to transcribe from scratch. Double that timeline now.

Consider all of the above when using this service. Do you have the time and resources to train your staff to know how to edit auto-generated captions and get them up to the appropriate standards? How challenging may your content be for the AI? Whenever and however you make the choice, make sure you deliver the best possible experience to your entire audience.

Closed captioning is an essential aspect of modern media consumption, bridging the gap of accessibility and inclusivity for diverse audiences. Yet, despite its widespread use, misconceptions about closed captioning persist. In this article, we delve into the most prevalent myths surrounding this invaluable feature, shedding light on the truth behind closed captioning's capabilities, impact, and indispensable role in enhancing the way we interact with video content.

Let’s debunk these common misunderstandings about closed captioning and gain a fresh perspective on the far-reaching importance of closed captioning in today's digital landscape.

Closed captioning is only for the deaf and hard of hearing

While closed captions are crucial for people with hearing impairments, they benefit a much broader audience. They are also helpful for people learning a new language, those in noisy environments, individuals with attention or cognitive challenges, and viewers who prefer to watch videos silently.

Closed captioning is automatic and 100% accurate

While there are automatic captioning tools, they are not always accurate, especially with complex content, background noise, or accents. Human involvement is often necessary to ensure high-quality and accurate captions.

Captions are always displayed at the same place on the screen

Some formats, like SCC, support positioning and allow captions to appear in different locations. However, most platforms use standard positioning at the bottom of the screen.

Captions can be added to videos as a separate file later

While it's possible to add closed captions after video production, it's more efficient and cost-effective to incorporate captioning during the production process. Integrating captions during editing ensures a seamless viewing experience.

Captions are only available for movies and TV shows

Closed captioning is essential for television and films, but it's also used in various other video content, including online videos, educational videos, social media clips, webinars, and live streams.

Captioning is a one-size-fits-all solution

Different platforms and devices may have varying requirements for closed caption formats and display styles. To ensure accessibility and optimal viewing experience, captions may need adjustments based on the target platform.

All countries have the same closed captioning standards

Captioning standards and regulations vary between countries, and it's essential to comply with the specific accessibility laws and guidelines of the target audience's location.

Closed captioning is expensive and time-consuming

While manual captioning can be time-consuming, there are cost-effective solutions available, including automatic captioning and professional captioning services. Moreover, the benefits of accessibility and broader audience reach often outweigh the investment.

In summary, closed captioning is a vital tool for enhancing accessibility and user experience in videos. Understanding the realities of closed captioning helps ensure that content creators and distributors make informed decisions to improve inclusivity and reach a broader audience.