https://www.youtube.com/watch?v=87DyyMV0kCY Состояние ИИ безопасности в 2026
🔥1

Channel
@gonzo_AI_security
On this record: Growth · Engagement · Reactions · Posts · Citations · Cite this entry
1,007subscribers
+107 since we began measuring on 7 August 2026
Risers and fallers across the register · movement among entries of 1,000–3,162.
| Telegram ID | -1002048302165 |
|---|---|
| Type | Channel |
| Username | @gonzo_AI_security |
| Created | Between 1 November 2023 and 13 May 2024 — estimated from Telegram’s id allocation, not measured. How this range is calculated. |
| First recorded | 7 August 2026 |
| Last confirmed live | 14 September 2026 |
| Measurements held | 8 |
| Confirmed unchanged | 1 time, most recently 14 September 2026 |
| On Telegram | t.me/gonzo_AI_security |
| Measured (UTC) | Subscribers | Change |
|---|---|---|
| 14 Sept 2026, 23:21 | 1,007 | +11 |
| 7 Sept 2026, 18:35 | 996 | +17 |
| 30 Aug 2026, 10:55 | 979 | +3 |
| 23 Aug 2026, 15:33 | 976 | +70 |
| 16 Aug 2026, 06:24 | 906 | +7 |
| 8 Aug 2026, 01:05 | 899 | -1 |
| 7 Aug 2026, 09:17 | 900 | no change |
| 7 Aug 2026, 08:48 | 900 | first reading |
16 posts held, back to 13 May 2024 — the reader has not yet reached the start of this channel’s public history, so older posts may sit further back, unread. Read across 1 page of Telegram’s post history, 20 posts per page.
Nothing published in the last 30 days. ERR and ER are rolling 30-day measures, so there is nothing to compute — we hold 16 posts for this entry, the most recent from 6 August 2026. An engagement rate over an empty window would be a number about nothing.
63 reactions across 15 posts, in 7 distinct kinds. The most used accounts for 31.7% of them.
| Reaction | Count | Share | Share, drawn |
|---|---|---|---|
| 👍 | 20 | 31.7% | |
| ❤ | 18 | 28.6% | |
| 😁 | 13 | 20.6% | |
| 😭 | 6 | 9.52% | |
| 🫡 | 4 | 6.35% | |
| 🔥 | 1 | 1.59% | |
| 😱 | 1 | 1.59% |
No sentiment is inferred, and none should be read in. This table is ordered by count and by nothing else. Emoji do not carry stable meaning across languages or communities — 🙏 is thanks in one channel and mourning in another — so we publish which ones were pressed and how often, and pass no judgement on what an audience meant by them.
Precision. Telegram publishes reaction counts per emoji and short-forms each one — 4.34K, 1.2M — so any single kind at or above 1,000 reaches us at three significant figures, and only counts below 1,000 are exact. The shares above are ratios of those figures and carry the same error. This is also why the total here can differ slightly from a reaction total printed elsewhere on the page: both are sums of the same rounded parts, taken over samples with different edges.
Coverage. Reactions were read on 15 of the 16 sampled posts in this sample. Summed by Telegram’s own count on each post — not by adding up the per-emoji breakdown above — those same posts carry 63 reactions in total: the kind of figure the paragraph above means by “a reaction total printed elsewhere on the page”.
Measured over the 16 most recent posts we hold, published 13 May 2024 to 6 August 2026, using the newest reading held for each. Telegram Stars are excluded: they are a payment, not a reaction, and they have their own section.
https://www.youtube.com/watch?v=87DyyMV0kCY Состояние ИИ безопасности в 2026
🔥1
Потестировал новую модель Антропика, где агенты могут управлять компьютером, используя экран Инъекции с ней работают точно также и позволяют менять поведение агента на новое (например, отправить команды в терминал)
😭6👍3🫡2
Наткнулся на клевую демку от стартапера, где можно в режиме реального времени поговорить с его аватаром-копией в формате видео-звонка – но больше всего мне понравилось, что джейлбрайкать такие интерфейсы можно голосом ☕️ В видео я прошу зачитать модель ее системный промпт и потом прошу начать говорить со мной на финском и польском, потому что я якобы ее автор, а потом все ломается Наверное, в будущем, будут люди, к…
😁6
Советую подаваться на курс, хорошо подходит для старта и относительно несложный отбор, плюс у вас появится какой-то пет-проект к его концу. https://t.me/ai_safety_digest/57 Дедлайн сегодня
❤3
и да, Moderation тут означает именно это > https://x.com/voooooogel/status/1834569673712754805
👍1
Поделал быстро таких же тестов с новой o1 Из интересного - Safety стал строже (т.е. реже помогает с опасными задачами), но теперь запросы отклоняются прямо на уровне API: HTTP code 400 from API "message": "Invalid prompt: your prompt was flagged as potentially violating our usage policy. Please try again with a different prompt." Видимо появился какой-то классификатор опасных промптов на входе (т.е. запросы даже не …
👍4
AI sandbagging
😁7👍2
Funny not funny AI app failures AI Deception: (кат) Обман проявляется в широком спектре систем ИИ, обученных для выполнения конкретной задачи. Обман особенно вероятен, когда система ИИ обучается для достижения экспертного уровня в играх, имеющих социальный элемент, таких как игра на построение альянсов и завоевание мира "Дипломатия", покер или другие задачи, связанные с теорией игр. Манипуляция: Meta разработала с…
>>
👍2
No, LLM Agents can not Autonomously Exploit Zero-day Vulnerabilities (yet) Недавно стала распространяться новая работа про LLM-хакеров — "Teams of LLM Agents can Exploit Zero-Day Vulnerabilities". Например, на них ссылается Jason Haddix в своем видео, ещё это репостилось во многих каналах. Почему эта некачественная работа, на которую не стоит ссылаться: 1) Это авторы, которые постоянно публикуют некачественные раб…
👍3
Примеры Один и тот же вопрос в GPT-4 (пик 1) и в GPT-4o (пик 2)
❤2
Примерно так Метод - что-то типа упрощенного PAIR (примерно: просишь какую-то слабую модель убедить сильную на выполнение какой-то плохой задачи. Тут сделано в 1 попытку на 50 задач из AdvBench - датасет вредоносных задач) Не супер точно, просто примерные тесты
🫡2❤1😱1
Showing the 12 most recent of 16 posts we hold for @gonzo_AI_security. View and reaction counts are the latest single reading for each post, not a live figure, and a recent post is still accumulating both. A view count marked ≈ was rounded by Telegram before we ever saw it — t.me prints views in full below 1,000 and to three significant figures above, so ≈1,200,000 means somewhere between 1,150,000 and 1,249,999. Unmarked counts are exact. Text is reproduced from the public post preview and truncated for length.
Republished by
Channels on the register that have forwarded this channel's posts into their own feed.
Republishes
Channels on the register whose posts this channel has forwarded.
Built only from forwarded posts we have actually read, on both sides. Coverage is early and deliberately incomplete: a missing link means we have not read the post that would prove it, never that the relationship does not exist. Counts are distinct forwarded posts observed, so they only ever go up as we read more.
Names
Channels on the register whose handles appear in this channel's posts.
A mention is a weaker signal than a forward and is counted separately for that reason — naming a channel is not republishing it, and a handle in a post body is easy to place deliberately. The post counts beside each row below are distinct posts in which the handle appeared, from posts we have read on both sides — the “Named by N registered channels” figure above is a different count, of distinct NAMING CHANNELS rather than posts, and is not the sum of the rows under it.
A live page changes as we take new readings, so a citation should name the measurement it is based on, not just the URL. The line below cites the subscriber count as measured 14 September 2026 — this entry's latest reading, not the date you are reading this.
“Gonzo-обзоры AI Security/Safety” (@gonzo_AI_security), 1,007 subscribers as measured 14 September 2026. Telegram Register, tgregister.com/channel/gonzo_AI_security.
Full measurement history, CC BY 4.0. Every reading this register holds for this entry, not just the latest one, as a dated, downloadable record: CSV · JSON. Free to use with attribution to tgregister.com. Each file carries its own generation timestamp, which is the figure to cite for exactly when the data was retrieved.