feat: enhance routing rule checks with NFT chain parsing and bypass validation Проблема с медиа в телеграм при использовании MTProto прокси

Fixes #298
This commit is contained in:
Daniel Lavrushin 2026-08-11 21:08:05 +02:00 committed by Daniel Lavrushin
parent 271fdb0ea9
commit b4168bb319
4 changed files with 193 additions and 27 deletions

View file

@ -2,11 +2,12 @@
## [1.76.1] - 2026-08-11
- FIXED: **Telegram never got past loading, with an empty chat list, in the WebSocket bridge routing mode** - the bridge was pinned to a Cloudflare Worker and was never offered Telegram's own WebSocket edge, the route the MTProto proxy had been taking all along for data centers 2 and 4, which is why the proxy carried the same account that the bridge could not. A Worker relay stops forwarding after some 13 to 17 KB of a connection and then holds it open in silence, so the client waited out its receive timeout, reconnected, and fetched the same first few kilobytes of its initial sync again, for as long as it was left running. Measured over 81 sessions on a censored network: 73 through a Worker, not one past 17 KB; 8 through Telegram's edge, one of them 1.1 MB.
- FIXED: **A Cloudflare Worker that had stopped relaying was handed the next Telegram connection as well** - a Worker stops forwarding without closing the WebSocket, so nothing in the connection says it is dead. The relay had no error to act on and held the session for the full five-minute idle timeout, while the client gave up on its own schedule and reconnected straight back onto the same route.
- FIXED: **The WebSocket bridge relayed to a data center address of its own choosing rather than the one the client had dialled** - it took the address from a table holding one entry per data center, while Telegram hands clients many, so a client on `149.154.167.92` was carried to `.91`, and a media address resolved to a data center whose entry in that table is not the media one.
- FIXED: **Connections b4 opens itself went out with none of b4's own bypass applied** - the mark that keeps them clear of b4's transparent-proxy diversion is also the mark the packet engine puts on everything it reinjects, and the firewall accepts that mark so a reinjected packet is not queued a second time. A fail-open dial, made at the moment a destination is already in trouble, was the one that left the machine most exposed by it.
- CHANGED: **The Cloudflare Worker script b4 pointed at ignored the port the client had asked for** - it opened 443 whatever it was told, and Telegram reaches a data center on port 80 as well, so every session on that transport was handed to an endpoint that does not speak it and answered nothing. The script is published in b4's own [Telegram documentation](https://daniellavrushin.github.io/b4/docs/mtproto) with the parameter honoured; a Worker deployed from the older page has to be redeployed to pick it up.
- FIXED: **Telegram never got past loading, with an empty chat list, in the WebSocket bridge routing mode** - the bridge only ever went through a Cloudflare Worker, never Telegram's own route, the one the MTProto proxy had been taking all along.
- FIXED: **A Cloudflare Worker that had stopped passing traffic kept being picked for new Telegram connections** - a session on a dead one was held for five minutes before it was given up.
- FIXED: **The WebSocket bridge sent traffic to a different data center address than the device had asked for** - it kept one address per data center, while Telegram hands out many.
- FIXED: **Connections b4 opens for itself went out with none of its own bypass applied** - a fail-open connection, made when a destination is already in trouble, was the one it left most exposed.
- FIXED: **b4 put its routing rules back only once they had disappeared completely** - part of a rule set going missing counted as healthy.
- CHANGED: **The Cloudflare Worker script b4 pointed at ignored the port it was asked for** - Telegram reaches a data center on a second port too, and those connections were answered by nothing. A script that reads the port is in the [Telegram documentation](https://daniellavrushin.github.io/b4/docs/mtproto); a Worker set up from the older page has to be deployed again.
## [1.76.0] - 2026-08-10

View file

@ -2,11 +2,12 @@
## [1.76.1] - 2026-08-11
- ИСПРАВЛЕНО: **Telegram не проходил дальше загрузки, со списком чатов пустым, в режиме маршрутизации через WebSocket-мост** - мост был привязан к Cloudflare Worker, и собственный WebSocket-узел Telegram ему не предлагался вовсе, хотя MTProto-прокси всё это время ходил в дата-центры 2 и 4 именно через него, поэтому один и тот же аккаунт прокси вытягивал, а мост нет. Релей через Worker перестаёт передавать данные примерно после 13-17 КБ в пределах соединения и дальше держит его открытым в тишине, поэтому клиент досиживал свой тайм-аут приёма, переподключался и заново забирал те же первые килобайты начальной синхронизации - и так сколько его ни оставь. Замерено на 81 сессии в цензурируемой сети: 73 через Worker, ни одна не прошла дальше 17 КБ; 8 через узел Telegram, одна из них 1,1 МБ.
- ИСПРАВЛЕНО: **Cloudflare Worker, переставший передавать данные, получал и следующее соединение Telegram** - Worker прекращает передачу, не закрывая WebSocket, поэтому в самом соединении ничто не говорит, что оно мертво. Реле было не на чем оборваться, и оно держало сессию все пять минут тайм-аута простоя, пока клиент сдавался по своему расписанию и переподключался ровно на тот же маршрут.
- ИСПРАВЛЕНО: **WebSocket-мост шёл на адрес дата-центра, выбранный им самим, а не на тот, куда обращался клиент** - адрес брался из таблицы с одной записью на дата-центр, тогда как Telegram выдаёт клиентам много адресов, поэтому клиента с `149.154.167.92` уносило на `.91`, а медийный адрес приводил к дата-центру, чья запись в этой таблице не медийная.
- ИСПРАВЛЕНО: **Соединения, которые b4 открывает сам, уходили без единой применённой к ним меры обхода** - метка, которая уводит их от собственного перехвата b4, это же и метка, которую движок ставит на всё, что переотправляет, а файрвол такую метку пропускает, чтобы переотправленный пакет не попал в очередь второй раз. Больше всего это оголяло машину как раз на fail-open соединении, которое открывается ровно тогда, когда с адресом уже что-то не так.
- ИЗМЕНЕНО: **Скрипт Cloudflare Worker, на который ссылался b4, игнорировал запрошенный клиентом порт** - он открывал 443 в любом случае, а Telegram ходит в дата-центр и на 80, поэтому каждая сессия на этом транспорте попадала к узлу, который её не понимает, и не получала в ответ ничего. Скрипт с поддержкой параметра опубликован в [документации b4 по Telegram](https://daniellavrushin.github.io/b4/ru/docs/mtproto); Worker, развёрнутый по старой странице, нужно задеплоить заново.
- ИСПРАВЛЕНО: **Telegram не проходил дальше загрузки, с пустым списком чатов, в режиме маршрутизации через WebSocket-мост** - мост ходил только через Cloudflare Worker, а собственный маршрут Telegram, которым всё это время шёл MTProto-прокси, ему не предлагался.
- ИСПРАВЛЕНО: **Cloudflare Worker, переставший пропускать трафик, продолжал доставаться новым соединениям Telegram** - сессию на мёртвом держали пять минут, прежде чем бросить.
- ИСПРАВЛЕНО: **WebSocket-мост отправлял трафик не на тот адрес дата-центра, к которому обращалось устройство** - он держал один адрес на дата-центр, тогда как Telegram выдаёт много.
- ИСПРАВЛЕНО: **Соединения, которые b4 открывает для себя, уходили без единой применённой к ним меры обхода** - сильнее всего это оголяло fail-open соединение, которое открывается как раз тогда, когда с адресом уже что-то не так.
- ИСПРАВЛЕНО: **b4 возвращал свои маршрутные правила на место, только когда они пропадали целиком** - пропажа части правил считалась нормой.
- ИЗМЕНЕНО: **Скрипт Cloudflare Worker, на который ссылался b4, игнорировал запрошенный порт** - Telegram ходит в дата-центр и на второй порт, и такие соединения не получали в ответ ничего. Скрипт, который порт читает, есть в [документации по Telegram](https://daniellavrushin.github.io/b4/ru/docs/mtproto); Worker, развёрнутый по старой странице, нужно развернуть заново.
## [1.76.0] - 2026-08-10

View file

@ -563,55 +563,101 @@ func RoutingRulesPresent(cfg *config.Config) bool {
switch eng := be.(type) {
case *routeNftBackend:
return routeNftRulesPresent()
return routeNftRulesPresent(cfg)
case *routeIptBackend:
return routeIptRulesPresent(eng, cfg)
}
return true
}
func routeNftRulesPresent() bool {
// parseNftRouteChains scans `nft list table inet b4_route` into the set of
// chains that exist and, per chain, the marks it returns on. Kept separate from
// the command so the brace handling can be tested against real output: a set's
// element list closes with a brace too, and mistaking that for the end of a
// chain would lose the rules that follow.
func parseNftRouteChains(out string) (present map[string]bool, bypass map[string]map[uint32]bool) {
present = make(map[string]bool)
bypass = make(map[string]map[uint32]bool)
chain := ""
for _, line := range strings.Split(out, "\n") {
line = strings.TrimSpace(line)
switch {
case strings.HasPrefix(line, "chain "):
chain = strings.TrimSpace(strings.TrimSuffix(line[len("chain "):], "{"))
present[chain] = true
continue
case line == "}":
chain = ""
continue
case chain == "":
continue
}
if m, verb, ok := nftParseMarkRule(line); ok && verb == "return" {
if bypass[chain] == nil {
bypass[chain] = make(map[uint32]bool)
}
bypass[chain][m] = true
}
}
return present, bypass
}
func routeNftRulesPresent(cfg *config.Config) bool {
out, err := run("nft", "list", "table", "inet", routeNftTable)
if err != nil || strings.TrimSpace(out) == "" {
return false
}
present := make(map[string]bool)
for _, line := range strings.Split(out, "\n") {
line = strings.TrimSpace(line)
if strings.HasPrefix(line, "chain ") {
name := strings.TrimSpace(strings.TrimSuffix(line[len("chain "):], "{"))
present[name] = true
}
}
present, bypass := parseNftRouteChains(out)
for _, st := range routeRuleCache {
for _, c := range routeStateChains(st) {
if !present[c.chain] {
return false
}
if !c.wantBypass {
continue
}
for _, m := range routeBypassMarks(cfg) {
if !bypass[c.chain][m] {
log.Tracef("Routing: chain %s lost its bypass on mark 0x%x", c.chain, m)
return false
}
}
}
}
return true
}
type routeChainRef struct{ chain, table string }
type routeChainRef struct {
chain, table string
wantBypass bool
}
func routeStateChains(st routeState) []routeChainRef {
switch {
case config.RoutingIsBlock(st.mode):
return []routeChainRef{{st.chainPre, "filter"}}
return []routeChainRef{{st.chainPre, "filter", false}}
case config.RoutingUsesTProxy(st.mode):
refs := []routeChainRef{{st.chainPre, "mangle"}}
refs := []routeChainRef{{st.chainPre, "mangle", true}}
if st.quicReject && st.chainQUIC != "" {
refs = append(refs, routeChainRef{st.chainQUIC, "filter"})
refs = append(refs, routeChainRef{st.chainQUIC, "filter", true})
}
return refs
default:
return []routeChainRef{{st.chainPre, "mangle"}, {st.chainOut, "mangle"}, {st.chainSNAT, "nat"}}
return []routeChainRef{
{st.chainPre, "mangle", true},
{st.chainOut, "mangle", true},
{st.chainSNAT, "nat", false},
}
}
}
// routeBypassMarks lists the marks a diverting chain must return on.
func routeBypassMarks(cfg *config.Config) []uint32 {
return []uint32{routeQueueBypassMark(cfg), SelfDialMark}
}
func routeIptRulesPresent(be *routeIptBackend, cfg *config.Config) bool {
needed := make(map[string]map[string]bool)
for _, st := range routeRuleCache {
@ -619,7 +665,7 @@ func routeIptRulesPresent(be *routeIptBackend, cfg *config.Config) bool {
if needed[c.table] == nil {
needed[c.table] = make(map[string]bool)
}
needed[c.table][c.chain] = true
needed[c.table][c.chain] = needed[c.table][c.chain] || c.wantBypass
}
}
if len(needed) == 0 {
@ -651,10 +697,23 @@ func routeIptRulesPresent(be *routeIptBackend, cfg *config.Config) bool {
}
}
}
for chain := range wantChains {
for chain, wantBypass := range wantChains {
if !present[chain] {
return false
}
if !wantBypass {
continue
}
spec, serr := run(cmd, "-w", "-t", table, "-S", chain)
if serr != nil {
continue
}
for _, m := range routeBypassMarks(cfg) {
if !strings.Contains(spec, fmt.Sprintf("--mark 0x%x/0x%x", m, m)) {
log.Tracef("Routing: chain %s lost its bypass on mark 0x%x", chain, m)
return false
}
}
}
}
}

View file

@ -1,6 +1,7 @@
package tables
import (
"strings"
"testing"
"github.com/daniellavrushin/b4/config"
@ -69,3 +70,107 @@ func TestRouteSelfDialBypass_EmitsBothMarks(t *testing.T) {
}
}
}
// Real `nft list table inet b4_route` output from a box running the bridge. The
// element list of a set closes with a brace on a content line, and the set block
// closes with one of its own, so a scanner that treats every brace as the end of
// a chain loses the rules that come after the sets.
const nftRouteTableSample = `table inet b4_route {
set b4r_3a97e38161af453_bf3d_v4 {
type ipv4_addr
flags interval,timeout
auto-merge
elements = { 91.105.192.0/23, 91.108.4.0-91.108.23.255,
91.108.56.0/22, 95.161.64.0/20,
149.154.160.0/20, 185.76.151.0/24 }
}
chain output {
type route hook output priority mangle - 1; policy accept;
meta mark & 0x00040000 == 0x00040000 return
meta mark & 0x00008000 == 0x00008000 return
ip protocol tcp ip daddr @b4r_3a97e38161af453_bf3d_v4 meta mark set 0x00024c9e
}
chain b4r_3a97e38161af453_bf3d_pre {
meta mark & 0x00008000 == 0x00008000 return
meta mark & 0x00040000 == 0x00040000 return
ip protocol tcp ip daddr @b4r_3a97e38161af453_bf3d_v4 meta mark set 0x00024c9e tproxy ip to :13686 accept
}
chain b4r_deadbeefdeadbee_0001_pre {
ip protocol tcp ip daddr @b4r_deadbeefdeadbee_0001_v4 drop
}
}`
func TestParseNftRouteChains(t *testing.T) {
present, bypass := parseNftRouteChains(nftRouteTableSample)
for _, c := range []string{"output", "b4r_3a97e38161af453_bf3d_pre", "b4r_deadbeefdeadbee_0001_pre"} {
if !present[c] {
t.Errorf("chain %s not found; the set block above it likely swallowed the scan", c)
}
}
for _, c := range []string{"output", "b4r_3a97e38161af453_bf3d_pre"} {
for _, m := range []uint32{0x8000, SelfDialMark} {
if !bypass[c][m] {
t.Errorf("chain %s: bypass on mark 0x%x not seen", c, m)
}
}
}
if len(bypass["b4r_deadbeefdeadbee_0001_pre"]) != 0 {
t.Error("a chain with no bypass rules must report none")
}
}
func TestParseNftRouteChains_MissingSelfDialBypass(t *testing.T) {
stripped := strings.ReplaceAll(nftRouteTableSample, "\t\tmeta mark & 0x00040000 == 0x00040000 return\n", "")
_, bypass := parseNftRouteChains(stripped)
if bypass["b4r_3a97e38161af453_bf3d_pre"][SelfDialMark] {
t.Fatal("a chain that lost its self-dial bypass must not report it as present")
}
if !bypass["b4r_3a97e38161af453_bf3d_pre"][0x8000] {
t.Error("the queue-mark bypass should still be seen")
}
}
func TestRouteStateChains_OnlyDivertingChainsWantBypass(t *testing.T) {
st := routeState{
mode: config.RoutingModeMTProtoWS,
chainPre: "pre",
chainOut: "out",
chainSNAT: "snat",
chainQUIC: "quic",
quicReject: true,
}
for _, c := range routeStateChains(st) {
if !c.wantBypass {
t.Errorf("tproxy mode chain %s diverts traffic and must be checked for its bypass rules", c.chain)
}
}
block := routeState{mode: config.RoutingModeBlock, chainPre: "pre"}
for _, c := range routeStateChains(block) {
if c.wantBypass {
t.Errorf("block chain %s carries no bypass rules; requiring them would make the monitor resync forever", c.chain)
}
}
iface := routeState{mode: config.RoutingModeInterface, chainPre: "pre", chainOut: "out", chainSNAT: "snat"}
want := map[string]bool{"pre": true, "out": true, "snat": false}
for _, c := range routeStateChains(iface) {
if c.wantBypass != want[c.chain] {
t.Errorf("interface chain %s: wantBypass=%v, expected %v", c.chain, c.wantBypass, want[c.chain])
}
}
}
func TestRouteBypassMarks_Distinct(t *testing.T) {
marks := routeBypassMarks(&config.Config{})
if len(marks) != 2 {
t.Fatalf("expected the queue mark and the self-dial mark, got %d", len(marks))
}
if marks[0] == marks[1] {
t.Errorf("both bypass marks are 0x%x", marks[0])
}
}