# My V8 Pi-hole instance with (Debian 12 / Proxmox VMs)

**URL:** <https://discourse.pi-hole.net/t/my-v8-pi-hole-instance-with-debian-12-proxmox-vms/85655>\
**Category:** Customizing Pi-hole\
**Created:** [April 6, 2026, 3:08am UTC](https://discourse.pi-hole.net/t/my-v8-pi-hole-instance-with-debian-12-proxmox-vms/85655 "2026-04-06T03:08:26Z")\
**Posts on this page:** 1\
**Showing post:** 7

<div class="post-metadata">

**Author:** ![smokingwheels](https://discourse-cdn.pi-hole.net/user_avatar/discourse.pi-hole.net/smokingwheels/32/2029_2.png) [@smokingwheels](https://discourse.pi-hole.net/u/smokingwheels)\
**Post date:** [April 8, 2026, 3:49am UTC](https://discourse.pi-hole.net/t/my-v8-pi-hole-instance-with-debian-12-proxmox-vms/85655/7 "2026-04-08T03:49:26Z")

</div>

## 🧪 DNS Performance Testing Summary (DNSdist + Pi-hole cluster)

I’ve been benchmarking a local DNS stack using **DNSdist (frontend)** with multiple **Pi-hole backends** under high-QPS UDP load (dnspyre/dnsblast style testing).

### 🖥 Hardware

- HP DL360 (older server)

- Proxmox VM environment

- Up to 24 cores available

* * *

## 🔧 Test Setup

- dnsdist as frontend load balancer

- Pi-hole instances as backends (scaled from a few → up to 12)

- High PPS / UDP-heavy workload (small packets, burst traffic)

- Both cached and uncached scenarios tested

👉 **NOTE: The results below primarily reflect _cached DNS performance_, not full recursive (uncached) resolution.**

* * *

## 📊 Key Findings

### 1. Peak vs Sustained Performance

- **Peak (short burst):** ~140k QPS

- **Sustained (realistic):** ~60k–70k QPS

➡ System handles very high bursts but settles to a stable throughput ceiling.

* * *

### 2. Frontend Core Scaling

- Increasing dnsdist cores alone did **not always improve performance**

- With limited backend capacity:

- With larger backend (12 Pi-holes)

➡ **Frontend scaling only helps if backend can absorb it**

* * *

### 3. Backend Scaling (most important factor)

- Adding more Pi-hole instances gave the biggest improvement

- Example:

➡ **Backend fanout had more impact than CPU tuning**

* * *

### 4. RAM Disk (tmpfs) Testing

- Moving logs/DB to RAM:

- But:

➡ Removing I/O bottlenecks can **increase overload effects**

* * *

### 5. Network Tuning (sysctl)

- Increased buffers/backlog improved burst handling

- Allowed ingestion of very high initial QPS (~400k+)

- But:

➡ **Higher buffers = more queueing, not more capacity**

* * *

### 6. Real Bottleneck

The limiting factors are NOT:

- CPU (low load observed)

- Disk (after tuning)

- Network stack (after sysctl tuning)

The bottleneck is:

- dnsdist processing + scheduling

- Pi-hole/FTL backend capacity

- Queue buildup under burst load

* * *

## 📉 Typical Stable Envelope

Across multiple runs:

- **Throughput:** ~65k–70k QPS

- **Error rate:** ~1.5%–3%

- **Latency (under burst):**

* * *

## 🧠 Key Takeaways

- **More cores ≠ more performance**

- **Backend scaling \> frontend scaling**

- **Burst capacity ≠ sustainable throughput**

- **Reducing bottlenecks can expose deeper limits**

- **Queueing (not CPU) drives latency at high load**

* * *

## 🏁 Best Performing Setup

So far:

- dnsdist: ~20 cores

- Backend: 12 Pi-hole instances

- No RAM disk tricks

- Tuned network stack

➡ Best balance of throughput and stability

* * *

## 💬 Final Thoughts

For high-QPS DNS workloads:

- Focus on **backend scaling and distribution**

- Avoid over-driving the system with unrealistic burst loads

- Tune for **sustained throughput** , not just peak numbers

* * *

## 🧪 Pi-hole (Tuned) Performance Snapshot Cached data

After tuning Pi-hole (logging adjustments, general cleanup, and backend balancing), I ran a focused test to look at individual node behaviour.

### 📊 Results (single-node style load 1000 domain list)

- **Achieved send QPS:** ~4,588

- **Total queries:** 67,164

- **Successful:** 66,484

- **Failed:** 680

- **Error rate:** ~ **1.01%**

- **Latency:**

* * *

### 📊 Results (single-node style load 20 domain list)

- **Achieved send QPS:** ~12k

## 🧠 Interpretation

- This reflects **per-node performance under sustained load** , not aggregate cluster throughput

- Results are consistent with earlier findings of **~12k cached peak per node** , dropping under sustained pressure

- Latency remains significantly lower than full-cluster burst tests due to:

* * *

## 🔍 Key Observations

- Pi-hole handles moderate sustained load well, but:

- Error rate (~1%) is acceptable for this test shape and aligns with earlier cluster behaviour

* * *

## 🏁 Takeaway

👉 Individual Pi-hole instances perform reliably in the **low-thousands QPS range under sustained load**  
👉 Scaling to higher throughput requires **horizontal backend expansion (more nodes)** rather than pushing individual instances harder

If anyone else is pushing Pi-hole/dnsdist at high PPS, would be interested to compare results 👍

---

_[View the full topic](https://discourse.pi-hole.net/t/my-v8-pi-hole-instance-with-debian-12-proxmox-vms/85655)._
