<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Mebularts Dev Blog]]></title><description><![CDATA[Mebularts Dev Blog]]></description><link>https://mebularts.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Mebularts Dev Blog</title><link>https://mebularts.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 07 Oct 2026 13:50:03 GMT</lastBuildDate><atom:link href="https://mebularts.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Designing a Reliable Proxy Rotation Strategy in Python]]></title><description><![CDATA[A proxy pool sounds simple until you actually have to keep one alive.
At first, the architecture usually looks like this:
request
   |
pick proxy
   |
send request
   |
done

Then production happens.
]]></description><link>https://mebularts.hashnode.dev/designing-a-reliable-proxy-rotation-strategy-in-python</link><guid isPermaLink="true">https://mebularts.hashnode.dev/designing-a-reliable-proxy-rotation-strategy-in-python</guid><category><![CDATA[Python]]></category><category><![CDATA[web scraping]]></category><category><![CDATA[automation]]></category><category><![CDATA[Backend Development]]></category><category><![CDATA[data-engineering]]></category><category><![CDATA[proxy]]></category><dc:creator><![CDATA[Mehmet Bulat]]></dc:creator><pubDate>Sat, 03 Oct 2026 13:56:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ac105e1c1e9ef319af35440/1ab7c739-cc9b-4b99-be66-7efe8e2f5938.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A proxy pool sounds simple until you actually have to keep one alive.</p>
<p>At first, the architecture usually looks like this:</p>
<pre><code class="language-text">request
   |
pick proxy
   |
send request
   |
done
</code></pre>
<p>Then production happens.</p>
<blockquote>
<p><a href="https://dataimpulse.com/?aff=1f1db2a8-e8f8-4dbd-bfd7-e95347b85b99">DataImpulse provides usage-based proxy options here →</a></p>
</blockquote>
<p>Some proxies time out. Some become extremely slow. Some work for five minutes and disappear. Others technically respond but fail repeatedly against the workload you actually care about.</p>
<p>At that point, random proxy selection is no longer enough.</p>
<p>A useful proxy layer needs to answer several questions:</p>
<ul>
<li><p>Which proxies are healthy?</p>
</li>
<li><p>Which ones are currently failing?</p>
</li>
<li><p>How long should a failed proxy stay out of rotation?</p>
</li>
<li><p>Should every request use a different IP?</p>
</li>
<li><p>When should a session keep the same IP?</p>
</li>
<li><p>How should latency affect proxy selection?</p>
</li>
<li><p>When does maintaining the pool cost more than using managed infrastructure?</p>
</li>
</ul>
<p>This article builds a small proxy rotation model around those questions.</p>
<blockquote>
<p>If the project already needs a managed residential, datacenter, or mobile proxy pool instead of maintaining public IPs manually, <a href="https://dataimpulse.com/?aff=1f1db2a8-e8f8-4dbd-bfd7-e95347b85b99">DataImpulse provides usage-based proxy options here →</a></p>
<p><em>Affiliate disclosure: that is an affiliate link. I may receive a commission from qualifying registrations or purchases at no additional cost to you.</em></p>
</blockquote>
<p>Now let’s build the logic behind the decision instead of blindly rotating IPs.</p>
<hr />
<h2>Random Rotation Is Not Enough</h2>
<p>The simplest implementation selects a random proxy:</p>
<pre><code class="language-python">import random

proxies = [
    "http://proxy-a:8080",
    "http://proxy-b:8080",
    "http://proxy-c:8080",
]

proxy = random.choice(proxies)
</code></pre>
<p>That works for a demo.</p>
<p>But suppose:</p>
<pre><code class="language-text">proxy-a -&gt; 95% success rate
proxy-b -&gt; 20% success rate
proxy-c -&gt; 80% success rate
</code></pre>
<p>Pure random selection still sends roughly one third of the traffic through <code>proxy-b</code>.</p>
<p>The pool contains quality information, but the algorithm ignores it.</p>
<p>A better solution is to track proxy health.</p>
<hr />
<h2>Representing Proxy Health</h2>
<p>Start with a small state model:</p>
<pre><code class="language-python">from dataclasses import dataclass, field
from time import time


@dataclass
class ProxyState:
    url: str

    requests: int = 0
    successes: int = 0
    failures: int = 0

    average_latency: float | None = None

    consecutive_failures: int = 0

    disabled_until: float = 0.0

    last_success: float | None = None
    last_failure: float | None = None
</code></pre>
<p>Each proxy now carries historical information rather than simply being a string.</p>
<p>We can calculate success rate:</p>
<pre><code class="language-python">def success_rate(proxy: ProxyState) -&gt; float:
    if proxy.requests == 0:
        return 1.0

    return proxy.successes / proxy.requests
</code></pre>
<p>For example:</p>
<pre><code class="language-text">Proxy A
requests: 100
successes: 94

success_rate = 0.94
</code></pre>
<p>Already much more useful than:</p>
<pre><code class="language-text">working = true
</code></pre>
<hr />
<h2>Latency Matters Too</h2>
<p>A proxy that succeeds consistently but takes 12 seconds per request may still be a poor choice.</p>
<p>We can maintain a rolling latency average.</p>
<pre><code class="language-python">def update_latency(proxy: ProxyState, latency: float):
    if proxy.average_latency is None:
        proxy.average_latency = latency
        return

    alpha = 0.25

    proxy.average_latency = (
        alpha * latency
        + (1 - alpha) * proxy.average_latency
    )
</code></pre>
<p>This uses an exponential moving average.</p>
<p>Recent latency has more influence, while older observations still contribute.</p>
<p>That is useful because network conditions change.</p>
<hr />
<h2>Creating a Health Score</h2>
<p>Now combine success rate and latency.</p>
<p>There is no universal scoring function, but a simple one is enough to demonstrate the idea:</p>
<pre><code class="language-python">def proxy_score(proxy: ProxyState) -&gt; float:
    reliability = success_rate(proxy)

    latency = proxy.average_latency or 1.0

    latency_penalty = min(
        latency / 10,
        0.5
    )

    failure_penalty = min(
        proxy.consecutive_failures * 0.1,
        0.5
    )

    return max(
        0,
        reliability
        - latency_penalty
        - failure_penalty
    )
</code></pre>
<p>Now proxies can be ranked.</p>
<pre><code class="language-python">ranked = sorted(
    proxies,
    key=proxy_score,
    reverse=True,
)
</code></pre>
<p>A healthy proxy naturally receives more traffic.</p>
<p>An unstable one gradually falls out of favor.</p>
<hr />
<h2>Do Not Retry a Dead Proxy Forever</h2>
<p>Another common mistake is retrying failed proxies immediately.</p>
<p>Imagine a proxy is completely offline.</p>
<p>Without cooldown logic:</p>
<pre><code class="language-text">Request 1 -&gt; Proxy A -&gt; failure
Request 2 -&gt; Proxy A -&gt; failure
Request 3 -&gt; Proxy A -&gt; failure
Request 4 -&gt; Proxy A -&gt; failure
</code></pre>
<p>That wastes time.</p>
<p>Instead, apply temporary quarantine.</p>
<pre><code class="language-python">from time import time


def mark_failure(proxy: ProxyState):
    proxy.requests += 1
    proxy.failures += 1
    proxy.consecutive_failures += 1
    proxy.last_failure = time()

    if proxy.consecutive_failures &gt;= 3:
        proxy.disabled_until = time() + 60
</code></pre>
<p>Then filter unavailable proxies:</p>
<pre><code class="language-python">def available(proxy: ProxyState) -&gt; bool:
    return time() &gt;= proxy.disabled_until
</code></pre>
<p>Now three consecutive failures remove the proxy from rotation for one minute.</p>
<hr />
<h2>Exponential Backoff Is Better</h2>
<p>A fixed cooldown works, but repeated failures should probably lead to longer cooldown periods.</p>
<pre><code class="language-python">def calculate_backoff(failures: int) -&gt; int:
    return min(
        2 ** failures,
        300
    )
</code></pre>
<p>Example:</p>
<pre><code class="language-text">Failure 1 -&gt; 2 seconds
Failure 2 -&gt; 4 seconds
Failure 3 -&gt; 8 seconds
Failure 4 -&gt; 16 seconds
Failure 5 -&gt; 32 seconds
...
Maximum -&gt; 300 seconds
</code></pre>
<p>Then:</p>
<pre><code class="language-python">proxy.disabled_until = (
    time()
    + calculate_backoff(
        proxy.consecutive_failures
    )
)
</code></pre>
<p>This prevents unstable proxies from continuously consuming resources.</p>
<hr />
<h2>Successful Requests Should Recover the Proxy</h2>
<p>A proxy should also be able to recover.</p>
<pre><code class="language-python">def mark_success(
    proxy: ProxyState,
    latency: float,
):
    proxy.requests += 1
    proxy.successes += 1

    proxy.consecutive_failures = 0
    proxy.last_success = time()

    update_latency(proxy, latency)
</code></pre>
<p>This is important because temporary network failures happen.</p>
<p>One bad request should not permanently remove a proxy.</p>
<hr />
<h2>Weighted Proxy Selection</h2>
<p>Instead of selecting the highest-scoring proxy every time, use weighted randomness.</p>
<p>Why?</p>
<p>Because always selecting the best proxy creates another problem:</p>
<pre><code class="language-text">best proxy
best proxy
best proxy
best proxy
best proxy
</code></pre>
<p>Eventually the “best” proxy receives almost all the load.</p>
<p>Weighted selection distributes traffic while still favoring healthier proxies.</p>
<pre><code class="language-python">import random


def choose_proxy(proxies):
    candidates = [
        proxy
        for proxy in proxies
        if available(proxy)
    ]

    if not candidates:
        raise RuntimeError(
            "No healthy proxies available"
        )

    weights = [
        max(proxy_score(p), 0.01)
        for p in candidates
    ]

    return random.choices(
        candidates,
        weights=weights,
        k=1,
    )[0]
</code></pre>
<p>Now a healthy proxy gets a higher probability rather than exclusive ownership of the workload.</p>
<hr />
<h2>Putting It Together</h2>
<p>A request function could look like this:</p>
<pre><code class="language-python">import time
import requests


def request_with_proxy(
    url,
    proxies,
    timeout=8,
):
    proxy = choose_proxy(proxies)

    started = time.perf_counter()

    try:
        response = requests.get(
            url,
            proxies={
                "http": proxy.url,
                "https": proxy.url,
            },
            timeout=timeout,
        )

        response.raise_for_status()

        latency = (
            time.perf_counter()
            - started
        )

        mark_success(
            proxy,
            latency,
        )

        return response

    except requests.RequestException:
        mark_failure(proxy)
        raise
</code></pre>
<p>The proxy pool is now adaptive.</p>
<p>Poor proxies slowly receive less traffic.</p>
<p>Repeated failures temporarily remove them.</p>
<p>Successful requests allow them to recover.</p>
<hr />
<h2>Rotation and Sessions Are Different Problems</h2>
<p>Many people use “proxy rotation” as if changing the IP on every request is always desirable.</p>
<p>It is not.</p>
<p>Consider this workflow:</p>
<pre><code class="language-text">GET /login
POST /login
GET /dashboard
GET /account
</code></pre>
<p>If every request comes from a different IP and possibly a different country, the application may behave differently.</p>
<p>Sometimes the desired behavior is:</p>
<pre><code class="language-text">Session 1
  Request 1 -&gt; IP A
  Request 2 -&gt; IP A
  Request 3 -&gt; IP A
  Request 4 -&gt; IP A
</code></pre>
<p>Then a new session can use:</p>
<pre><code class="language-text">Session 2 -&gt; IP B
</code></pre>
<p>This is the idea behind a <strong>sticky session</strong>.</p>
<p>Rotation and stickiness are not opposites in the sense that one is better.</p>
<p>They solve different requirements.</p>
<hr />
<h2>A Simple Sticky Assignment</h2>
<p>For an internal pool, a mapping can associate an application session with a proxy:</p>
<pre><code class="language-python">session_proxy = {}


def proxy_for_session(
    session_id,
    proxies,
):
    current = session_proxy.get(
        session_id
    )

    if (
        current
        and available(current)
    ):
        return current

    selected = choose_proxy(proxies)

    session_proxy[session_id] = (
        selected
    )

    return selected
</code></pre>
<p>Now the session remains consistent unless the proxy becomes unhealthy.</p>
<hr />
<h2>Public Proxy Pools Have a Maintenance Cost</h2>
<p>This architecture is useful, but it also exposes an important fact.</p>
<p>The bigger a public proxy pool becomes, the more infrastructure you need around it.</p>
<p>You start adding:</p>
<ul>
<li><p>Validation workers</p>
</li>
<li><p>Health scoring</p>
</li>
<li><p>Persistent storage</p>
</li>
<li><p>Geo-IP verification</p>
</li>
<li><p>Retry queues</p>
</li>
<li><p>Cooldowns</p>
</li>
<li><p>Monitoring</p>
</li>
<li><p>Metrics</p>
</li>
<li><p>Alerting</p>
</li>
<li><p>Re-testing jobs</p>
</li>
<li><p>Proxy replacement</p>
</li>
</ul>
<p>At some point, the question changes.</p>
<p>It stops being:</p>
<blockquote>
<p>How can I build a better proxy checker?</p>
</blockquote>
<p>and becomes:</p>
<blockquote>
<p>Why am I operating proxy infrastructure at all?</p>
</blockquote>
<hr />
<h2>When I Would Use Datacenter Proxies</h2>
<p>Datacenter proxies are usually the first managed option I would test.</p>
<p>They often provide:</p>
<ul>
<li><p>High throughput</p>
</li>
<li><p>Low latency</p>
</li>
<li><p>Lower pricing</p>
</li>
<li><p>Predictable infrastructure</p>
</li>
</ul>
<p>If the target workload works with datacenter IPs, there may be little reason to immediately move to residential traffic.</p>
<p>This is also why comparing providers purely by residential proxy pricing can be misleading.</p>
<p>The cheapest proxy is often the cheapest <strong>appropriate</strong> proxy type.</p>
<hr />
<h2>When Residential Proxies Become Useful</h2>
<p>Residential proxies become more relevant when the project requires things such as:</p>
<ul>
<li><p>ISP-associated IP addresses</p>
</li>
<li><p>Geographic diversity</p>
</li>
<li><p>Country-specific testing</p>
</li>
<li><p>Local search monitoring</p>
</li>
<li><p>International pricing research</p>
</li>
<li><p>Large rotating IP pools</p>
</li>
</ul>
<p>A managed network also removes much of the health-checking infrastructure described above.</p>
<p>Instead of maintaining thousands of individual IPs, the application connects to a gateway and lets the provider handle the pool.</p>
<hr />
<h2>An Example Managed Option</h2>
<p>One provider that exposes residential, datacenter, and mobile proxy products under one platform is DataImpulse.</p>
<p>At the time of writing, its published entry pricing includes:</p>
<ul>
<li><p><strong>Residential:</strong> from $1/GB</p>
</li>
<li><p><strong>Datacenter:</strong> from $0.50/GB</p>
</li>
<li><p><strong>Mobile:</strong> from $2/GB</p>
</li>
</ul>
<p>Its residential network is advertised as including more than <strong>90M residential IPs across 195+ countries</strong>.</p>
<p>That makes the architecture different.</p>
<p>Instead of:</p>
<pre><code class="language-text">Application
    |
local proxy database
    |
health checker
    |
rotation service
    |
thousands of public proxies
</code></pre>
<p>you can potentially move toward:</p>
<pre><code class="language-text">Application
    |
managed proxy gateway
    |
provider-managed IP pool
</code></pre>
<p>For workloads where maintaining the pool has become its own subsystem, that trade-off can be worth evaluating.</p>
<p><a href="https://dataimpulse.com/?aff=1f1db2a8-e8f8-4dbd-bfd7-e95347b85b99">Explore current DataImpulse residential, datacenter and mobile proxy options →</a></p>
<hr />
<h2>Testing Public Proxies Is Still Useful</h2>
<p>None of this means public proxies are useless.</p>
<p>They are great for:</p>
<ul>
<li><p>Learning</p>
</li>
<li><p>Experiments</p>
</li>
<li><p>Networking projects</p>
</li>
<li><p>Proxy checker development</p>
</li>
<li><p>Proof-of-concept systems</p>
</li>
</ul>
<p>I also maintain an independent open-source proxy testing project:</p>
<p><a href="https://github.com/mebularts/proxy_tester">Proxy Tester on GitHub →</a></p>
<p>This project is completely separate from DataImpulse.</p>
<p>The public proxies tested with it are <strong>not</strong> DataImpulse IPs, and the project does not use DataImpulse's residential, datacenter, or mobile proxy pools.</p>
<hr />
<h2>A Better Architecture</h2>
<p>For a production-oriented proxy layer, I would separate responsibilities:</p>
<pre><code class="language-text">                   +------------------+
                   |  Application     |
                   +--------+---------+
                            |
                            v
                   +------------------+
                   | Proxy Selector   |
                   +--------+---------+
                            |
             +--------------+--------------+
             |                             |
             v                             v
     +---------------+             +---------------+
     | Health Store  |             | Session Store |
     +---------------+             +---------------+
             |
             v
     +---------------+
     | Proxy Pool    |
     +---------------+
</code></pre>
<p>The selector should not know everything.</p>
<p>It should ask other components for:</p>
<ul>
<li><p>Current health</p>
</li>
<li><p>Session assignment</p>
</li>
<li><p>Availability</p>
</li>
<li><p>Proxy capabilities</p>
</li>
</ul>
<p>This makes it much easier to replace the underlying proxy source later.</p>
<p>Public list today.</p>
<p>Datacenter provider tomorrow.</p>
<p>Residential pool later.</p>
<p>The application itself does not have to care.</p>
<hr />
<h2>The Real Metric: Successful Work</h2>
<p>The more proxy infrastructure I build, the less interested I become in nominal proxy price.</p>
<p>Suppose one provider costs:</p>
<pre><code class="language-text">$0.50 / GB
</code></pre>
<p>but generates a high retry rate.</p>
<p>Another costs:</p>
<pre><code class="language-text">$1.00 / GB
</code></pre>
<p>but successfully completes almost every operation.</p>
<p>The first is cheaper only if the application actually gets useful work from that traffic.</p>
<p>The metric I care about is closer to:</p>
<pre><code class="language-text">total infrastructure cost
-------------------------
successful operations
</code></pre>
<p>And infrastructure cost includes more than bandwidth:</p>
<pre><code class="language-text">proxy traffic
+ retries
+ compute
+ engineering time
+ monitoring
+ maintenance
</code></pre>
<p>That calculation often changes which solution is actually “cheap.”</p>
<hr />
<h2>Final Thoughts</h2>
<p>A reliable proxy system is not simply:</p>
<pre><code class="language-python">random.choice(proxy_list)
</code></pre>
<p>A production-oriented pool benefits from:</p>
<ul>
<li><p>Health tracking</p>
</li>
<li><p>Latency measurements</p>
</li>
<li><p>Weighted selection</p>
</li>
<li><p>Failure cooldowns</p>
</li>
<li><p>Exponential backoff</p>
</li>
<li><p>Recovery logic</p>
</li>
<li><p>Sticky session support</p>
</li>
<li><p>Metrics</p>
</li>
</ul>
<p>Building those systems is useful engineering work.</p>
<p>It is also exactly what reveals when operating a public proxy pool is no longer worth the effort.</p>
<p>For experiments, public proxies can be excellent.</p>
<p>For sustained workloads, managed datacenter or residential proxy infrastructure may reduce the operational burden significantly.</p>
<p>The important part is to optimize for <strong>successful work</strong>, not simply the lowest advertised proxy price.</p>
<hr />
<h3>Disclosure</h3>
<p>This article contains affiliate links to DataImpulse. If a qualifying registration or purchase is made through those links, I may receive a commission at no additional cost to the reader.</p>
<p>The GitHub Proxy Tester linked above is my independent open-source project and is not affiliated with or supplied by DataImpulse.</p>
<p>Use proxy infrastructure and automated data collection only where you have the necessary authorization and in accordance with applicable law and service terms.</p>
]]></content:encoded></item></channel></rss>