https://www.reddit.com/r/homeassistant/comments/1vhofep/hacked_and_debloated_an_echo_dot_2_local_llm/

Channel
Speech Technology
@speechtech
On this record: Growth · Engagement · Posts · Citations · Cite this entry
1,711subscribers
+26 since we began measuring on 6 August 2026
Risers and fallers across the register · movement among entries of 1,000–3,162.
Register entry
| Telegram ID | -1001472248479 |
|---|---|
| Type | Channel |
| Username | @speechtech |
| Created | 2 April 2020 — measured — cross-checked against a third-party dataset (TGDataset) |
| First recorded | 6 August 2026 |
| Last confirmed live | 3 September 2026 |
| Measurements held | 10 |
| Confirmed unchanged | 1 time, most recently 3 September 2026 |
| On Telegram | t.me/speechtech |
Growth
| Measured (UTC) | Subscribers | Change |
|---|---|---|
| 3 Sept 2026, 15:17 | 1,711 | +2 |
| 31 Aug 2026, 03:27 | 1,709 | +6 |
| 28 Aug 2026, 02:34 | 1,703 | +6 |
| 25 Aug 2026, 10:55 | 1,697 | +5 |
| 22 Aug 2026, 18:04 | 1,692 | -3 |
| 19 Aug 2026, 07:24 | 1,695 | +3 |
| 16 Aug 2026, 08:38 | 1,692 | +3 |
| 10 Aug 2026, 02:41 | 1,689 | +6 |
| 6 Aug 2026, 21:20 | 1,683 | -2 |
| 6 Aug 2026, 13:51 | 1,685 | first reading |
Engagement
22 posts held, back to 13 July 2026 — the reader has not yet reached the start of this channel’s public history, so older posts may sit further back, unread. Read across 2 pages of Telegram’s post history, 20 posts per page.
Nothing published in the last 30 days. ERR and ER are rolling 30-day measures, so there is nothing to compute — we hold 22 posts for this entry, the most recent from 7 August 2026. An engagement rate over an empty window would be a number about nothing.
Recent posts
There is a big interest in full duplex as I see, here is a nice collection of papers https://github.com/Ruiqi-Yan/Awesome-Full-Duplex-SDM
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B NVIDIA NemotronLabs VoiceChat is a 11B end-to-end, real-time speech full duplex (FD) model for conversational AI that jointly performs streaming speech understanding and speech generation [1, 2]. Unlike traditional cascaded stacks (ASR → LLM → TTS), this model achieves full duplex, real-time, seamless voice interaction in one unified architecture, elimi…
Things move on in OpenAI as well. Interesting that voice model is separate. And no turn detector anymore. https://x.com/OpenAI/status/2084378415818579975 Lots of interesting technical details, from realtime inference, to dynamic compaction, to WebRTC optimization. https://openai.com/index/continuous-voice-interaction-with-gpt-live/
One more https://github.com/AnXMuy/AgenticASR
https://huckiyang.github.io/voice-memory/ from NVIDIA https://arxiv.org/abs/2607.26410 Voice Memory for Agentic Speech Recognition Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain this http URL and decides per utterance whether…
https://huggingface.co/nyralabs/CrisperWhisper2.0_large Most speech-to-text systems never actually decide whether to write down what was said or what was meant. They inherit that choice from their training data and apply it inconsistently. CrisperWhisper 2.0 makes it an explicit, controllable choice. One recording, two transcripts: Verbatim, exactly what was said, in one consistent format: [um] so we we need to, to …
Some recent Uzbek things https://huggingface.co/datasets/k2speech/FeruzaSpeech - single speaker 40 hours TTS dataset https://huggingface.co/collections/navai-uz/navai-whisper-collection - recently trained Whisper models from Navai https://navai.pro https://huggingface.co/instinct-org/collections - some loosely organized data https://huggingface.co/datasets/OvozifyLabs/asr_evaluate_set - evaluation dataset with Te…
Everyone builds self-improvement loops in LLMs, I wonder how they could look like in ASR/TTS. Not many publications on that yet.
We compared three LALM judges against a calibrated human panel across 15 dimensions of speech quality. The LALMs tracked humans closely on relevance, answer quality, and instruction following—what was said—but were much less reliable on naturalness, emotion, pronunciation, and overall feel—how it was said. https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl
Interesting math on speech LLM https://arxiv.org/abs/2604.08003v1 Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs Yuan Xie, Jiaqi Song, Guang Qiu, Xianliang Wang, Ming Lei, Jie Gao, Jie Wu Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promisin…
Interesting project https://github.com/Xiaobin-Rong/unipase UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Xiaobin Rong, Zheng Wang, Yushi Wang, Jun Gao, Jing Lu Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tail…
Showing the 12 most recent of 22 posts we hold for @speechtech. View and reaction counts are the latest single reading for each post, not a live figure, and a recent post is still accumulating both. A view count marked ≈ was rounded by Telegram before we ever saw it — t.me prints views in full below 1,000 and to three significant figures above, so ≈1,200,000 means somewhere between 1,150,000 and 1,249,999. Unmarked counts are exact. Text is reproduced from the public post preview and truncated for length.
Forward network
Republished by
Channels on the register that have forwarded this channel's posts into their own feed.
Built only from forwarded posts we have actually read, on both sides. Coverage is early and deliberately incomplete: a missing link means we have not read the post that would prove it, never that the relationship does not exist. Counts are distinct forwarded posts observed, so they only ever go up as we read more.
Cite this entry
A live page changes as we take new readings, so a citation should name the measurement it is based on, not just the URL. The line below cites the subscriber count as measured 3 September 2026 — this entry's latest reading, not the date you are reading this.
“Speech Technology” (@speechtech), 1,711 subscribers as measured 3 September 2026. Telegram Register, tgregister.com/channel/speechtech.
Full measurement history, CC BY 4.0. Every reading this register holds for this entry, not just the latest one, as a dated, downloadable record: CSV · JSON. Free to use with attribution to tgregister.com. Each file carries its own generation timestamp, which is the figure to cite for exactly when the data was retrieved.