Re-examining "Connection prematurely closed by remote server" with Unbound 1.17.1

Re-examining "Connection prematurely closed by remote server" with Unbound 1.17.1

I wanted to follow up on the earlier investigation in this thread because I found some additional information in Unbound's source that seems relevant to the behavior originally observed here.

I have also requested that Unbound issue #1237 be reconsidered/reopened based on this additional evidence.

Related reports

Original Pi-hole investigation:

Related Unbound issue #1237:

In the original sync.opera.com investigation, DL6ER showed that Pi-hole received both A and HTTPS queries, forwarded them to Unbound, and that Unbound had already obtained the information needed to answer the A query. However, the A request ultimately ended with no waiting replies and no downstream response, while Pi-hole reported:

failed to send TCP(read_write) packet:
Connection prematurely closed by remote server

One question raised during that investigation was:

“Maybe dnsmasq can be configured to send multiple queries over a single TCP connection?”

I think that question may be more significant than it initially appeared.

Unbound 1.17.1 tracks multiple requests on a TCP stream

In the Unbound 1.17.1 source, TCP connections have a tcp_req_info structure that tracks multiple open and completed requests:

int num_open_req;
struct tcp_req_open_item* open_req_list;

int num_done_req;
struct tcp_req_done_item* done_req_list;

Mesh states are associated with that TCP request state, so multiple requests can exist on the same TCP communication point.

The discard-timeout path in 1.17.1 is not UDP-only

In 1.17.1, the discard_timeout check can call:

comm_point_drop_reply(&r->query_reply);

comm_point_drop_reply() then has explicit TCP handling:

if(repinfo->c->type == comm_udp)
    return;

if(repinfo->c->tcp_req_info)
    repinfo->c->tcp_req_info->is_drop = 1;

reclaim_tcp_handler(repinfo->c);

So the old code path is capable of marking the TCP request state as dropped and reclaiming the TCP handler.

That seems potentially important if multiple DNS requests are sharing a TCP connection.

Unbound later changed this behavior

Unbound 1.25.0 subsequently changed the discard-timeout handling so that it applies to UDP but not active stream connections. The release notes describe this as:

“Fix to drop UDP for discard-timeout, but not stream connections.”

The newer implementation explicitly checks for UDP before applying the old-reply discard logic, with a comment explaining that stream replies are not discarded because the stream remains open and the client is waiting for an answer.

I am not claiming that this later change definitively fixes the original Pi-hole issue. I don't think the available evidence is enough to establish that.

What I do think is worth investigating is whether the original behavior may involve the interaction between:

  • multiple requests on a single DNS-over-TCP connection,

  • Unbound's tcp_req_info handling,

  • the old discard-timeout implementation, and

  • connection-level TCP handler reclamation.

Why this still matters for Pi-hole users

This is particularly relevant because Debian 12 (Bookworm) continues to provide Unbound 1.17.1 through its normal package repositories.

Although Unbound 1.25.0+ exists upstream, moving from the distribution-provided 1.17.1 to a much newer upstream version is not a routine package upgrade for someone who needs to remain on Debian 12. It can mean using a package outside the distribution's normal versioning/support path or upgrading the underlying OS.

So if the later Unbound change is related to the behavior originally observed here, there are still users for whom simply upgrading to 1.25.0+ is not a practical solution.

Current testing

I've also been testing this locally with Pi-hole → Unbound over 127.0.0.1:5335.

A controlled synthetic workload using 600 TCP and 600 UDP queries against known-good domains did not reproduce the failure. I have nevertheless continued to see the original Connection prematurely closed by remote server error intermittently under normal use.

I also tested with:

discard-timeout: 0

That did not eliminate all of the errors, including separate Resource temporarily unavailable events, so I'm not suggesting that discard-timeout explains every instance of the problem.

What I'm interested in determining

The main questions I think are now:

  1. Does Pi-hole/dnsmasq ever issue multiple DNS requests over the same TCP connection in the affected scenario?

  2. Could the 1.17.1 Unbound behavior described above cause one or more requests on that stream to be lost when the TCP handler is reclaimed?

  3. Was the later change to make discard-timeout UDP-only intended to prevent this type of stream-handling behavior?

  4. If so, would a backport of the relevant Unbound change be an appropriate solution for users who remain on 1.17.x?

I have posted the source-level findings to Unbound issue #1237 as well, since that issue tracks the same Connection prematurely closed by remote server symptom and a very similar Raspberry Pi/Debian 12/Unbound 1.17.1 environment.

I'm mainly posting this here because the original investigation was conducted from the Pi-hole side, and the unanswered question about multiple queries sharing a TCP connection may be relevant to understanding what is actually happening between FTL/dnsmasq and Unbound.