Telegram RegisterThe public register of Telegram
Telegram profile photo for Тестирование и оценка ИИ

Channel

Тестирование и оценка ИИ

@testingofai

On this record: Growth · Engagement · Reactions · Posts · Citations · Cite this entry

1,362subscribers

+2 since we began measuring on 7 August 2026

Risers and fallers across the register · movement among entries of 1,000–3,162.

Register entry

Telegram ID-1002518310614
TypeChannel
Username@testingofai
CreatedBetween 1 March 2025 and 31 July 2025 — estimated from Telegram’s id allocation, not measured. How this range is calculated.
First recorded7 August 2026
Last confirmed live7 September 2026
Measurements held9
Confirmed unchanged1 time, most recently 7 September 2026
On Telegramt.me/testingofai

Growth

1,3601,3641,3627 August 2026 — 1,360 subscribers7 August 2026 — 1,360 subscribers14 August 2026 — 1,361 subscribers20 August 2026 — 1,360 subscribers24 August 2026 — 1,363 subscribers27 August 2026 — 1,361 subscribers30 August 2026 — 1,364 subscribers2 September 2026 — 1,363 subscribers7 September 2026 — 1,362 subscribers7 August 20267 September 2026
9 measurements spanning 31 days, net +2. Dots are measurements; the straight line between them is drawn to join them, not to claim we know the path taken in between — snapshots are recorded only when a count changes, so gaps mean “no change observed”, never “interpolated”. The vertical axis spans 1,359–1,365 and does not start at zero.
Measurement log — every subscribers count we have recorded
Measured (UTC)SubscribersChange
7 Sept 2026, 07:361,362-1
2 Sept 2026, 15:031,363-1
30 Aug 2026, 21:341,364+3
27 Aug 2026, 23:341,361-2
24 Aug 2026, 14:191,363+3
20 Aug 2026, 14:341,360-1
14 Aug 2026, 00:081,361+1
7 Aug 2026, 07:471,360no change
7 Aug 2026, 00:301,360first reading

Engagement

20 posts held, back to 22 May 2026the reader has not yet reached the start of this channel’s public history, so older posts may sit further back, unread. Read across 1 page of Telegram’s post history, 20 posts per page.

Nothing published in the last 30 days. ERR and ER are rolling 30-day measures, so there is nothing to compute — we hold 20 posts for this entry, the most recent from 4 August 2026. An engagement rate over an empty window would be a number about nothing.

Reaction mix

149 reactions across 20 posts, in 5 distinct kinds. The most used accounts for 43.6% of them.

Every reaction kind recorded on the sample, most used first
ReactionCountShareShare, drawn
🔥6543.6%
👍3926.2%
2818.8%
💯128.05%
🤔53.36%

No sentiment is inferred, and none should be read in. This table is ordered by count and by nothing else. Emoji do not carry stable meaning across languages or communities — 🙏 is thanks in one channel and mourning in another — so we publish which ones were pressed and how often, and pass no judgement on what an audience meant by them.

Precision. Telegram publishes reaction counts per emoji and short-forms each one — 4.34K, 1.2M — so any single kind at or above 1,000 reaches us at three significant figures, and only counts below 1,000 are exact. The shares above are ratios of those figures and carry the same error. This is also why the total here can differ slightly from a reaction total printed elsewhere on the page: both are sums of the same rounded parts, taken over samples with different edges.

Coverage. Reactions were read on 20 of the 20 sampled posts in this sample. Summed by Telegram’s own count on each post — not by adding up the per-emoji breakdown above — those same posts carry 149 reactions in total: the kind of figure the paragraph above means by “a reaction total printed elsewhere on the page”.

Measured over the 20 most recent posts we hold, published 22 May 2026 to 4 August 2026, using the newest reading held for each. Telegram Stars are excluded: they are a payment, not a reaction, and they have their own section.

Recent posts

4 Aug 2026, 10:59 UTC201 views9 reactionsread 7 August 2026
Photo

Хочу поделиться апдейтом, который я недавно добавил в свою библиотеку eval-ai-library, и заодно рассказать, почему я вообще над этим работал. Когда команды берут готовые инструменты оценки, такие как DeepEval, Ragas, Promptfoo, то они получают набор стандартных метрик: faithfulness, answer relevancy, contextual precision и так далее. Работает это отлично, пока ваша AI система решает более-менее типовую задачу. RAG н

🔥81

30 Jul 2026, 14:23 UTC334 views4 reactionsread 7 August 2026
Photo

Пока весь фокус на capability-бенчмарках, буквально пару недель назад восесь стран (Сингапур, Япония, Австралия, Канада, Франция, Кения, Корея, Великобритания) + ЕС провели совместное тестирование безопасности агентных ИИ. Они прогнали ИИ агентов через примерно 1500 задач и 1200 тулов на девяти языках, от английского до кисуахили и главный результат, который они получили, что safety pass rate у лучшей модели состав

👍4

28 Jul 2026, 10:49 UTC296 views4 reactionsread 7 August 2026
Photo

Сегодня поговорим о Robustness testing через призму метаморфического тестирования. Немного разберем, что такое метаморфическое тестирование. Это метод тестирования программного обеспечения, при котором проверка корректности основана не на самих результатах, а на метаморфических отношениях - ожидаемых зависимостях между входами и выходами программы. Проще говоря: Мы задаём системе вход X и получаем результат Y. Зате

👍21🔥1

24 Jul 2026, 07:38 UTC348 views6 reactionsread 7 August 2026
Photo

Недавно изучил исследование "Coin Flip Judge?" , в котором исследователи прогнали 29 задач из 10 категорий через две judge-модели от OpenAI, повторяя одну и ту же пару ответов много раз. Результат получился интересным, потому что судья меняет вердикт в среднем в 13,6% прогонов. На 28% вопросов отклонение было выше 20%, а на одном вопросе судья менял решение в 56% случаев, то есть по сути как подбрасывать монетку. П

👍4🤔2

23 Jul 2026, 07:41 UTC368 views7 reactionsread 7 August 2026
Photo

Всем привет! У меня отличные новости, я определился с датой старта 4-го потока на курсе по оценке и тестированию ИИ систем! Ииии…. мы начинаем 9 сентября! Если вы давно хотели изучить для себя эту область тестирования, но откладывали, то сейчас до 19 августа вы можете записаться на курс и получить скидку 10% на все обучение. Для записи нужно просто оставить заявку на сайте: eval-ai.com Напомню, что это едиственны

🔥52

14 Jul 2026, 13:39 UTC463 views8 reactionsread 7 August 2026
Photo

Кажется то, что происходит с бенчмарками агентов последние пару месяцев, реально стоит обсудить. В апреле Berkeley RDI взломали восемь индустриальных agent-бенчмарков, таких как Terminal-Bench, SWE-bench, WebArena, OSWorld, GAIA и другие, и получили результаты, близкие к 100%, вообще не решая задачу по существу. Где-то модель просто подделывала pass результат. Где-то агент скачивал эталонный файл по ссылке из конфи

👍3🤔3🔥2

9 Jul 2026, 12:28 UTC464 views11 reactionsread 7 August 2026
Photo

Недавно подумал насколько мы привыкли использовать ИИ в своей работе, особенно кодинг агентов, типа Cursor или Claude Code, что начали постепенно забывать по аспекту связанные с качеством и что немаловажно эффективность работы такого агента. И я пришел к этому выводу не случайно. Буквально недавно я создал мультиагента ИИ агента на базе claude code sdk, который решает различные задачи связанные с тестирование, такие

💯65

7 Jul 2026, 07:15 UTC405 views9 reactionsread 7 August 2026
Photo

Сегодня разберем частые ошибки, которые могут встречаться в в работе AI агентов. Если тестируете AI агентов, то этот список поможет сфокусироваться на ключевых проблемах. 1. Зацикливание (Infinite Loops). Агент повторяет одни и те же действия бесконечно, например, агент пытается получить информацию, но получает ошибку или считает, что информации недостаточно для продолжения работы и повторяет запрос снова и снова.

👍3💯32🔥1

1 Jul 2026, 07:25 UTC477 views7 reactionsread 7 August 2026
Photo

Для всех, кто использует ИИ, но всегда задавался вопросом, а как же собственно LLM работают, я хочу поделиться подборкой видео от 3blue1brown. 3Blue1Brown объясняют сложные концепции через визуализации, которые делают абстрактные идеи понятными. В нескольких видео они разбирают принципы и архитектуру работы LLM, показывая, как LLM работает под капотом. Что вы узнаете из видео: Tokenization - как текст превращается

4🔥2👍1

29 Jun 2026, 12:45 UTC453 views10 reactionsread 7 August 2026
File

Всем привет! Сегодня хочу поделиться своим новым исследованием, которое прошло рецензирование и было опубликовано в Journal of Electrical Systems and Information Technology. Ссылка на публикацию В статье проводится анализ существующих подходов к оценке качества систем, построенных на основе генеративного искусственного интеллекта. Рассматриваются лексические методы (TF-IDF и BM25), семантические эмбеддинги, гибрид

🔥6👍4

24 Jun 2026, 15:29 UTC558 views10 reactionsread 7 August 2026
Photo

AI Harness в тестировании В индустрии сейчас все больше набирает популярность новый модный термин - agent harness. Это весь софт вокруг LLM, который превращает “модель, отвечающую на промпт” в агента, который реально делает работу. Инструменты, управление контекстом, цикл “действие → результат → следующий шаг”, память, обработка ошибок, все это является harness. Почему это важно? Сегодня в том же SWE-bench одна и т

👍4🔥41💯1

17 Jun 2026, 07:07 UTC540 views4 reactionsread 7 August 2026
Photo

Сегодня поговорим о том, почему оценка качества AI-систем стала критически важной. Современный AI переживает революцию: от простых классификаторов мы дошли до моделей, которые пишут код, создают изображения, управляют роботами. Но вместе с этим пришли и новые вызовы. 1. Масштаб. Модели обучаются на терабайтах данных, и классические методы вроде cross-validation уже не работают, потому что это слишком дорого и сложн

🔥4

Showing the 12 most recent of 20 posts we hold for @testingofai. View and reaction counts are the latest single reading for each post, not a live figure, and a recent post is still accumulating both. A view count marked was rounded by Telegram before we ever saw it — t.me prints views in full below 1,000 and to three significant figures above, so ≈1,200,000 means somewhere between 1,150,000 and 1,249,999. Unmarked counts are exact. Text is reproduced from the public post preview and truncated for length.

Mentions

Named by 4 registered channels — every channel on the register whose own posts have named this one, by its current username or any other username it currently holds, merged from two separately captured readings of the same fact so a namer caught by only one of them is not missed and a namer both caught is not counted twice. A username this channel has since dropped is not matched — that handle may belong to someone else now, and crediting today’s namer to yesterday’s owner would misattribute it.

Names

Channels on the register whose handles appear in this channel's posts.

A mention is a weaker signal than a forward and is counted separately for that reason — naming a channel is not republishing it, and a handle in a post body is easy to place deliberately. The post counts beside each row below are distinct posts in which the handle appeared, from posts we have read on both sides — the “Named by N registered channels” figure above is a different count, of distinct NAMING CHANNELS rather than posts, and is not the sum of the rows under it.

Cite this entry

A live page changes as we take new readings, so a citation should name the measurement it is based on, not just the URL. The line below cites the subscriber count as measured 7 September 2026 — this entry's latest reading, not the date you are reading this.

“Тестирование и оценка ИИ” (@testingofai), 1,362 subscribers as measured 7 September 2026. Telegram Register, tgregister.com/channel/testingofai.

Full measurement history, CC BY 4.0. Every reading this register holds for this entry, not just the latest one, as a dated, downloadable record: CSV · JSON. Free to use with attribution to tgregister.com. Each file carries its own generation timestamp, which is the figure to cite for exactly when the data was retrieved.