Download Latest Version v.202610.post2 source code.zip (26.0 MB) Google Add to Preferred Sources
Home / v.202610.post1
Name Modified Size InfoDownloads / Week
Parent folder
ESPnet version 202610.post1 source code.tar.gz 2026-09-20 21.3 MB
ESPnet version 202610.post1 source code.zip 2026-09-20 25.9 MB
README.md 2026-09-20 4.8 kB
Totals: 3 Items   47.2 MB 0

Summary

Overview

A patch on top of 202610, and nothing about installing or upgrading changes — same Python, same PyTorch, same extras. What changes is how much of ESPnet you can reach without writing a program.

espnet demo serves the OWSM browser demo on localhost in one command, and espnet asr --live transcribes the microphone as you speak — a window at a time, with a ceiling on what it holds, so a decoder slower than real time drops audio and says so rather than growing until the machine runs out. import espnet now works: espnet.load(tag) returns the inference object for a model without your knowing which of the twenty-five classes serves it. Three more demos run as Spaces — ljspeech-vits, universal-se, speaker-verification — maintained from this repository like the two OWSM ones. And espnet/espnet:inference-cpu-latest (0.6 GB) or -gpu-latest runs a published model with nothing installed at all.

The S2T interface is one class again. s2t_inference.Speech2Text loads either kind of OWSM checkpoint — encoder-decoder or CTC-only — and decodes four ways through two methods: __call__ for a search (attention, hybrid, or CTC prefix beam) and best_path() for CTC decoding with no search, which is what Speech2TextGreedySearch was. Both classes in s2t_inference_ctc are now deprecated forwarders, and everything in the repository calls the new one. If you use ctc_weight=1.0 expecting greedy decoding, it now says out loud that you are getting a prefix beam search instead: the two are both called "CTC decoding", and on one 20 s window measured here the search took 55 s against best_path's 16 s on OWSM-CTC, and 119 s against 0.9 s on OWSM v3.1 base.

Three correctness fixes: the Conv2dSubsamplingWOPosEnc mask is rebuilt from ilens rather than from a stale length (#6670), mask indexing and batching in partially autoregressive inference (#6673), and S2TPreprocessor had been dropping the data-augmentation arguments it was given (#6718).


Requirements — nothing to do

202610 202610.post1
Python >=3.12,<3.14 unchanged
PyTorch 2.11.0, 2.13.0, 2.14.0 unchanged
pip install espnet inference only; training is espnet[train] unchanged

A patch release, so upgrading is pip install -U espnet and nothing else.

espnet asr --live needs a microphone library that is deliberately not a dependency: pip install sounddevice, plus PortAudio from your system package manager. espnet demo needs pip install espnet[demo] for gradio, which is not a dependency either — no inference path should need a web framework to import.

Important PRs

Recipe

  • PR [#6732]: Follow espnet/notebook's new layout (by @sw005320)
  • PR [#6723]: Add TTS, speech enhancement and speaker verification demo Spaces (by @sw005320)
  • PR [#6722]: espnet demo: the OWSM browser demo as one command (by @sw005320)
  • PR [#6695]: docs: add Bagpiper vLLM inference guide (by @whr-a)

Bugfix

  • PR [#6718]: Two fixes the course notebook needed: hydra effects and S2T data augmentation (by @sw005320)
  • PR [#6673]: Fix mask indexing and batching in partially AR inference (by @meryemsakin)
  • PR [#6670]: fix(asr): rebuild Conv2dSubsamplingWOPosEnc mask from ilens (by @modelpath-dev)

Documentation

  • PR [#6744]: A patch release is .postN, which is what PyPI accepts (by @sw005320)
  • PR [#6721]: Add a small inference docker image, and a table of which image is which (by @sw005320)
  • PR [#6719]: Move the CI coverage detail out of the README (by @sw005320)
  • PR [#6716]: Say 202610 in What's new, and make the next release say its own (by @sw005320)

Refactoring

  • PR [#6728]: One Speech2Text for both kinds of S2T checkpoint (by @sw005320)

Others

  • PR [#6737]: Enhance with USES in the MCP server too (by @sw005320)
  • PR [#6734]: Run the inference images before pushing them (by @sw005320)
  • PR [#6725]: Transcribe as the audio arrives: espnet asr --live and --stream (by @sw005320)
  • PR [#6720]: Add a top-level espnet package: import espnet and espnet.load() (by @sw005320)
  • PR [#6717]: Give every human pull request the next release milestone (by @sw005320)
  • PR [#6715]: Answer espnet --version (by @sw005320)

Contributors

4 contributors, across 18 merged pull requests in v.202610.post1.

@meryemsakin, @modelpath-dev, @sw005320, @whr-a.

Full changelog: https://github.com/espnet/espnet/compare/v.202610...v.202610.post1

Source: README.md, updated 2026-09-20