← back to blog
June 30, 2026·9 min read

Parsing Telegram audience: where to find leads who actually reply

When someone sends me a «ready base of 50000 usernames in niche X», the first thing I ask is what percentage of those is active in the last 30 days. The answer is always vague or absent. That is the first sign the dump is useless.

In this post I will explain how we collect Telegram leads inside Sndwich, what filters actually cut the noise, and why a username without activity is just a 16-character string, not a lead.

Where Telegram leads come from

There are basically four sources, each with its quirks:

  1. Channel comment chats. Best source. If a person wrote a comment they are alive, not the admin, and the topic interests them. Downside: slower collection because Telegram returns messages in batches.
  2. Open group members. Decent source, lower quality. Many people joined a group long ago and are inactive now.
  3. Channel subscriber list. Available only to channel admins via the Bot API. If you are not an admin, forget about it.
  4. Post reactions. A relatively new source. Telegram returns who put a reaction, and that is a very warm audience. Not enabled on all channels and slower to parse.

How we parse inside Sndwich

Technically the process looks like this. You have a worker Telegram account. Through Telethon (a Python library for Telegram MTProto) we open the channel or group, walk the message history or the member list, and collect fields needed for downstream segmentation.

For each collected user:

  • tg_user_id (numeric)
  • username (may be empty)
  • first_name / last_name (may be empty)
  • date of last message in this channel
  • text of the last message (for context)
  • whether it was a comment, repost, or reaction

This is raw material. It has to be filtered, otherwise you will spam a pile of dead accounts and earn a restriction for a high rejection rate.

Filters that actually work

1. Activity in the last 30 days

The main filter. If a user has not posted in this channel for 30 days, the chance they open your DM is low. By default we drop everyone whose last_message_at is older than 30 days.

2. Has a username

Without a username you cannot reliably message the user. Telegram allows messaging by numeric id but that only works when the user has DMs from non-contacts enabled. Many do not, and you hit a wall.

3. Not a bot, not an admin

Bots show up in every dump. Easy to spot via the is_bot flag or by a username ending in «bot». Admins also reply to comments on their own channel, you want to drop them so you do not pitch the channel owner about their own niche.

4. Comment quality

The subtle filter. Most Telegram comments are «thanks», «cool», emojis, spam. Genuinely meaningful comments are 5-15% of the total volume. Those 5-15% are your warm leads. You can filter by comment length (30 characters is a decent threshold) and by presence of niche-specific words.

5. Language

If your outreach is in English, there is no point messaging people who comment in Spanish or Chinese. A simple language detector (cld3, fastText) takes seconds per thousand messages and saves your reply rate.

Ethical line: what NOT to collect

Legally, parsing public channels is a grey zone. Not illegal (the data is open) but not unambiguously fine under GDPR either. A few lines I would not cross:

  • Private / invite-only groups. If you were let in as a regular user and you parse the member list, that looks like breach of trust. Do not.
  • Full message history per author. Collect last_message for context, not a person's whole conversation. That edges into persistent profiling.
  • Geo and private info from bios. A person listing a phone number in their bio does not mean you can use it. Ignore those fields.

Volume and realistic numbers

Rough numbers from real channels so you have a sense of what to expect:

  • Channel with 50k subscribers, active, SaaS topic = ~3000-5000 commenters in 90 days. After filters, 200-400 usable leads.
  • Group with 10k members = ~1500-2500 ever posted. After filters, 100-200.
  • A news channel typically = 2-5% activity relative to total size.

So out of 50000 «all subscribers» you really work with 300-500. If someone promises 50000 usable leads from a single channel they either never ran outreach or they are lying. Probably the latter.

How Sndwich works under the hood

Two parsing modes:

Discovery

You type keywords in any of 19 languages, we search for relevant channels via contacts.Search in Telegram. You get back a list with metrics: subscriber count, activity percentage (from a message sample), language, whether comments are enabled, etc. You pick 5-10 interesting ones and move to the next step.

Channel parse

Picked channels parse in the background. The worker walks message history with 30-120 second pauses between batches (Telegram likes a human pace), collects commenters with the filters described above, and lands them in your leads table.

In a day, one worker account parses around 5-10 medium-activity channels. If you need more, add more workers (or use the shared pool, we automate that).

What to do with these leads

Parsing is not the goal in itself. The goal is to run a drip campaign that gets replies. That is the next post, how to write Telegram cold DMs that get replies, not blocks.

If you do not want to deal with the plumbing, try Sndwich. Discovery and channel parse work during the first 7 days of the trial, no card required.