Bot Management Demo — Cheat Sheet
~15–20 min · Glance once, click, talk · Deeper details in Bot_Management_REFERENCE.docx
Pre-demo setup checklist
- Run
warp-cli status in Terminal and confirm it says Disconnected, not just "paused." On a Cloudflare-managed laptop, WARP auto-reconnects on its own after a pause (org policy, Auto Connect setting) — checking the menu bar icon once isn't enough, it can silently come back before or during the demo. If it reconnects faster than you can finish, don't fight it: run the live tests and the backdrop generator from a phone on cellular, a personal unmanaged device, or a cloud VM instead of your work laptop.
- Why this matters: WARP-tunneled traffic from an enrolled device gets treated very differently by Bot Management than normal internet traffic. Confirmed live — the exact same curl/Python traffic that scored 99 "Likely human" while WARP was connected scored 1 "Automated" the moment it came from a genuinely unmanaged network. If Test 2 shows a high/human score right before you present, check WARP first before assuming the zone config broke.
- WAF Custom Rule deployed:
(http.user_agent eq "") and (http.host eq "nginx.tarheel.us") → Block
- Redirect Rule deployed:
(http.user_agent contains "python-requests") and (http.host eq "nginx.tarheel.us") → 302 to a learning page
- Access app
NGINX Landing has a Bypass / Everyone policy (or is deleted)
- Terminal open with big font
- Browser tab queued to
https://nginx.tarheel.us/
- 5-10 minutes before you start: run
python3 bot_demo_traffic_generator.py (same folder as this file). It sends a mix of honest bots, UA-lying bots, fake crawlers, a credential-stuffing burst pattern, and header-realistic "human-ish" traffic at nginx.tarheel.us, so Step 2's charts and JA4 cards have real shape instead of near-empty graphs. It stops on its own (default 3 min) — let it finish before you walk in.
1. Open (~2 min)
Set the table, then roll into the demo. If these questions come up on their own, let them.
"Let me set this up quickly. If you have a login page, a pricing page, or an API, bots are already hitting it. That part's a given. The real question is: can you tell the bots apart from your actual customers? Most teams can't. That's what we fix."
"And this is different from the WAF. The WAF looks for attacks — is this request trying to break something? We look at who's behind it — is this a person, or a machine? Because a request can look completely normal and still be a bot copying your prices all day long. Basic plans catch the obvious bots. The ones that get through are the sneaky ones — bots that look and act just like a real browser. Those are the ones we go after."
"So how do we spot them? The old way was to just believe what the visitor tells you. A browser says 'I'm Chrome,' and you take its word for it — but anyone can type that in. So we don't take their word for it. We give every visitor a score from 1 to 99, based on clues that are really hard to fake: things like their connection, their reputation, and how they behave. Here's the main point to remember: all those clues add up to one score. You make your decisions on the score. Let me show you."
2. Bot Analytics page (~4 min)
▶ CLICK: Security → Analytics → Bot analysis tab
a) Likely automated vs. Likely human counts
"Right at the top, traffic is split into 'Likely Human' and 'Likely Automated.' High-level view of every request that hit this site in the last 24 hours — already sorted so you can see how much is just background noise."
What's actually behind this graph. If you ran the backdrop generator before the demo, this isn't organic traffic, it's a deliberate mix I sent a few minutes ago: some bots being honest about what they are, some lying about being a browser, a few claiming to be Googlebot or Bingbot, a bursty credential-stuffing style pattern against a login path, and some traffic with perfect browser-style headers that still isn't a real browser underneath. That's why the split isn't 50/50 or silent — every one of these personas gets caught the same way, none of them can fake the underlying TLS handshake, only the User-Agent header, which is the whole point we're about to prove live with five hand-picked requests.
b) Detection sources — ML vs. Cloudflare service
"Check out the Bot Score Source. A huge chunk of scoring comes from a machine-learning model that's basically seen it all. Cloudflare sits in front of about 20% of the internet, so we spot a brand-new attack pattern the second it pops up anywhere — and protect your site instantly."
c) SWITCH TO EVENTS TAB — JA4 Fingerprint cards
"If you want the 'how' behind the magic, look at these JA4 fingerprints. JA4 is specifically the TLS fingerprint — a hash of how a visitor introduces itself during the encrypted handshake: the protocol version, the cipher list, the extensions, the signature algorithms. Think of it as the serial number of the software's network stack. Even if a bot rotates its IP or lies about its name in the User-Agent, the TLS handshake still looks like whatever actually made the connection."
"One thing I want to be straight about up front, because it's the most common misunderstanding: JA4 is a signal, not a switch. We don't just match a JA4 and block it. Millions of real people share the exact same JA4 as any given Chrome build — that's normal. What makes JA4 powerful is what we do with it: we feed it into the Bot Score, we compare it against the other layers to catch clients that are lying, and we track it across our whole network. I'll show you all three."
3. Live proof — five requests, five distinct signals (~6 min)
"I'm going to send five requests to my own site from a terminal. All five will return the same page — the demo isn't about blocking things in the terminal, it's about the signals Cloudflare records on each request. After the requests, we'll flip to Security Analytics and I'll show you how each one produced a visibly different row. Detection is the hard part. Once you can detect the difference between these five clients, writing a rule to block, challenge, or route any of them is a two-minute job in the WAF."
→ Switch to terminal. Big font on projector. Run all five in quick succession — under 30 seconds total — so they group together in the analytics timeline.
The Ray ID is your receipt. Every command uses -sI to show only response headers instead of the full HTML body. The line you care about is cf-ray — paste any of those Ray IDs into Security Analytics to see the exact row for that request, including bot score, JA4, verified_bot status, and detection IDs. That's how you prove Cloudflare made the judgment, not the origin.
What you should expect in the terminal. All five commands probably return 200 OK from your demo host unless you've configured a WAF or Bot Fight Mode rule that acts on bot score. That's expected — this demo shows Cloudflare's detection capability, which is the valuable half. Enforcement is a rule the customer writes after they see the signals. Don't promise blocks in the terminal that you haven't configured on the host.
TEST 1 — Real browser (baseline) → analytics signal: bot score in the human range, JA4 identifies Chrome
Open in Chrome: https://nginx.tarheel.us/?demo=t1
"Real browser. Real TLS handshake. Client Hints headers sent by Chrome automatically. In analytics this row should show a bot score in the human range, JA4 identifying Chrome, and the UA column matching the JA4. This is your 'what a human looks like' reference row. To grab the Ray ID from a real browser, open DevTools → Network tab → click the request → look at Response Headers for cf-ray."
TEST 2 — Honest curl (no deception) → analytics signal: low bot score, JA4 identifies curl
curl -sI "https://nginx.tarheel.us/?demo=t2" | grep -E "HTTP|cf-ray"
"Unmodified curl. User-Agent literally says curl/8.x. JA4 fingerprint says curl. In analytics: bot score in the bot range, JA4 identifies curl, UA matches JA4 — everything is consistent. This is the honest bot baseline. Grab the cf-ray value; you'll paste it into Analytics in Step 4."
TEST 3 — Lying curl (User-Agent says Chrome, TLS says curl) → analytics signal: UA/JA4 mismatch — same JA4 as Test 2
curl -sI -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" \
"https://nginx.tarheel.us/?demo=t3" | grep -E "HTTP|cf-ray"
"Same curl binary as Test 2. Same JA4. But the User-Agent now claims Chrome. This is the payoff row. In Analytics the UA column will say Chrome, the JA4 column will identify curl. Those two facts cannot both be true for a legitimate client — a real Chrome browser would produce a Chrome JA4. That's the deception, visible in one row. Same JA4 as Test 2 is the receipt: same client, one is being honest about what it is, one is lying."
Copy this JA4 down. When you open this row in Analytics, grab the value in the JA4 column — that exact string is what you'll paste into the mismatch rule in Step 5b (cf.bot_management.ja4 eq "<this JA4>" and http.user_agent contains "Chrome"). Tests 2 and 3 share the JA4, so either row gives you the same value; Test 3 is just the one that proves why the rule needs the User-Agent condition.
TEST 4 — Fake Googlebot → analytics signal: verified_bot = false, UA claims Googlebot
curl -sI -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
"https://nginx.tarheel.us/?demo=t4" | grep -E "HTTP|cf-ray"
"Curl is claiming to be Googlebot. In analytics, the User-Agent column will say Googlebot. The verified_bot column will say false. Cloudflare cross-checks Googlebot claims against Google's actual published IP ranges — this request is not from a Google IP, so it fails verification. A WAF that only reads User-Agent strings would treat this as Googlebot and let it through. Cloudflare labels it as an unverified claim, which is exactly the signal a customer needs to write a rule against. That rule is a two-minute job once you can see the signal."
TEST 5 — Empty User-Agent → analytics signal: blank UA column, distinct row shape
curl -sI -A "" "https://nginx.tarheel.us/?demo=t5" | grep -E "HTTP|cf-ray"
"Empty User-Agent. Almost no legitimate client sends this. In analytics, the UA column is literally blank — a visually distinct row shape that stands out against the other four. Look for a numeric detection ID here that didn't appear on the other tests; that's Cloudflare flagging the empty-UA condition specifically."
🛠️ Optional — Run all four curl tests at once and collect Ray IDs
Paste this whole block into your terminal. It runs Tests 2–5 and prints a labeled table of Ray IDs so you can walk to Analytics with the receipts in hand.
run_test() {
LABEL="$1"; UA="$2"; TAG="$3"
RAY=$(curl -sI -A "$UA" "https://nginx.tarheel.us/?demo=$TAG" | grep -i '^cf-ray:' | awk '{print $2}' | tr -d '\r')
printf "%-20s cf-ray=%s\n" "$LABEL" "$RAY"
}
run_test "Honest curl" "curl/8.7.1" t2
run_test "Lying curl (Chrome)" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" t3
run_test "Fake Googlebot" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" t4
run_test "Empty UA" "" t5
Output looks like:
Honest curl cf-ray=a23d54749c8912c5-ORD
Lying curl (Chrome) cf-ray=a23d54812fab12c5-ORD
Fake Googlebot cf-ray=a23d548a3b6f12c5-ORD
Empty UA cf-ray=a23d5490c78b12c5-ORD
PUNCHLINE:
"Five requests. Same URL, same 200 response, same body. But in the analytics dashboard, five visibly different rows — different bot scores, different verified-bot status, different UA-to-JA4 relationships, different detection IDs. That's the detection layer. Once you can see that difference, acting on any of these clients — challenge, block, or route — is a one-rule change in the WAF."
"The one that usually sells it: Test 2 and Test 3 have the same JA4 fingerprint. Same TLS stack, same client. But Test 3 says 'I am Chrome' in the User-Agent. That contradiction is the tell — the client can lie in headers, it cannot lie in the TLS handshake without rewriting its TLS library. And notice what the winning rule actually keys on: not 'block this JA4,' but 'block this JA4 when it's also lying about being a browser.' The fingerprint is the ingredient; the mismatch is the meal."
4. Walk the dashboard — five rows, five stories (~4 min)
"Let me show you what Cloudflare recorded for those five requests. Each row tells a different piece of the story."
▶ CLICK: Security → Analytics → Bot analysis tab.
▶ FILTER: Paste a cf-ray value from the terminal output into the Ray ID filter, or filter by the query string tag demo=t2 through t5. Sort by time to see all five in order.
What to point at, per row:
- Test 1 (Chrome): Bot score in the human range. JA4 identifies Chrome. User-Agent matches JA4. Client Hints populated. "This is what a human looks like."
- Test 2 (honest curl): Bot score in the bot range. JA4 identifies curl. UA says curl. Numeric detection IDs may appear in the detection column. "Honest bot — every signal agrees."
- Test 3 (lying curl): Same JA4 as Test 2. User-Agent now says Chrome. "UA and JA4 disagree — impossible for a real Chrome client. This is the deception, visible in one row."
- Test 4 (fake Googlebot): User-Agent says Googlebot.
verified_bot: false. "IP check failed — not actually Google. This is the row that separates Cloudflare from any WAF that only reads User-Agent strings."
- Test 5 (empty UA): User-Agent field is blank. Distinct row shape. "Nothing legitimate sends an empty User-Agent."
"The JA4 fingerprint is what makes this real. Tests 2 and 3 came from the same client — same binary, same TLS handshake — but Test 3 tried to lie in the header layer. That deception is visible in the dashboard as two columns that disagree. All of this was computed before any WAF rule ran. It's metadata on every request, available in Logpush, queryable in Analytics, usable in any custom rule."
Honest caveats on the numeric detection IDs. Cloudflare doesn't publish which detection ID number means what — that's deliberate policy so attackers can't reverse-engineer the heuristics. The story is that different rows fire different IDs, not that any specific ID has a named meaning. If a customer asks "what does detection ID 50331656 mean?", the honest answer is that mapping is available in an NDA session with a Bots product manager, arranged through their account team. Also, on a low-traffic demo host, some detections may not trip at all — pivot to the JA4 column and the verified_bot column, which will always show the story regardless of which specific detection IDs fired.
5. Act on the signal — the payoff (~4 min)
"Here's what most teams do today — a bot misbehaves, they look at the IP, they block the IP. Bot rotates IPs, they block more IPs. Whack-a-mole, forever. IPs are throwaway. So the instinct is: okay, block the JA4 instead. And that's almost right — but if you just hard-block a JA4, you'll take down real customers, because legitimate browsers share fingerprints. Let me show you the way that actually works."
a) The primary rule — act on the Bot Score
▶ CLICK: Security → WAF → Custom Rules → Create Rule
(cf.bot_management.score lt 20) → Managed Challenge
"This is the rule 90% of customers should lead with. The Bot Score already has JA4 baked into it, alongside IP reputation, behavior, and the machine-learning model. So instead of betting everything on one fingerprint, you're acting on the aggregate of every signal at once. Score under 20, Managed Challenge — a real human sails through invisibly, a headless bot gets stuck. No single point of failure, nothing to maintain as fingerprints drift."
b) The precision rule — catch the liar (JA4 mismatch)
(cf.bot_management.ja4 eq "<paste the JA4 you copied from the Test 3 row>"
and http.user_agent contains "Chrome") → Block
"This is where JA4 earns its keep. Remember Test 3 — the User-Agent said Chrome, the JA4 said curl. Paste that JA4 (the one you copied a moment ago) in here. The rule isn't 'block this JA4,' it's 'block this JA4 only when it's also claiming to be a browser it isn't.' A real curl user hitting your API is untouched — they don't claim to be Chrome, so this rule skips them. Only the client that's actively lying gets blocked. Compound logic, not a blunt block."
c) The emergency rule — targeted, temporary JA4 block
(cf.bot_management.ja4 eq "<attacker JA4>") → Block [remove after the attack]
"There is a time to hard-block a JA4 outright: an active attack. If you're being scraped or hit with an L7 flood right now, and your logs show thousands of requests across many IPs all sharing one non-browser JA4, block that JA4 and the attack stops instantly. But this is a fire extinguisher, not a smoke detector — you pull the rule once the attack subsides, because that same fingerprint could belong to a legitimate client next month as TLS libraries shift."
The one-liner to leave them with: IPs are license plates — trivial to swap. JA4 is the engine's serial number — much harder to change. But you don't impound every car with a given engine. You act on the score for everyday defense, you act on the mismatch to catch liars, and you hard-block a JA4 only to stop an attack in progress.
5b. If asked: "Then why does Bot Management even show me the JA4?" (~2 min)
"Fair question — if blanket-blocking it is a bad idea, why surface it at all? Three reasons, and none of them is 'so you can block it.'"
- Explainability. When Cloudflare scores a request a 1, your security team needs to know why. The JA4 in the log is the evidence: 'this claimed to be Chrome on Windows, but the fingerprint is a Python runtime.' It turns the machine-learning black box into something an engineer can audit.
- Signals Intelligence. Cloudflare tracks every JA4 across the whole network and publishes stats per fingerprint — things like
browser_ratio (what share of this JA4's traffic looks like real browsers), uas_rank (how many different User-Agents claim this one TLS signature), and IP diversity. If a JA4 shows a browser ratio near zero, that's your cue it's safe to build targeting logic around it. Those signals feed the score automatically, and you can also read them yourself.
- Compound rules. You saw this in Step 5 — the JA4 is the ingredient that lets you write 'block only when the fingerprint and the User-Agent disagree,' which is far more surgical than any score threshold alone.
"So we show it for visibility and precision, not for one-click blocking. The blunt instrument is IP blocking. JA4 is the scalpel — and a scalpel is only useful if you can see what you're cutting."
5c. If asked: "Can't a bot just change its JA4?" (~2 min)
"Short answer: not easily, and when they do, it usually hurts them more than it helps. This is the whole reason JA4 replaced the older JA3 standard."
- You can't randomize your way out. The old JA3 trick was to shuffle your cipher order to look different every time. JA4 sorts the ciphers and extensions before hashing, so that randomization does nothing.
- It's baked into the network stack. The JA4 comes from low-level TLS characteristics — protocol version, extensions, signature algorithms, ALPN. To change it, you have to change the actual TLS library your client uses, not flip a header. That means tools like curl-impersonate or a custom TLS stack, and if any one parameter is slightly off, the inconsistency itself becomes a detection signal.
- Rotating it constantly backfires. Real humans don't change their TLS parameters mid-session. A fingerprint that shifts every few requests is itself anomalous — it flags the session rather than hiding it. The bots that blend in best hold a fixed, realistic fingerprint, which is exactly the kind of thing our Signals Intelligence and cross-layer checks are built to catch.
"So the honest framing for the customer: JA4 isn't unbeatable, nothing is. But changing it is expensive, brittle, and often self-defeating — which is why it's one of the strongest single signals we have, and why it's most valuable as an input to the score rather than a standalone block."
If asked: "Why Cloudflare and not Akamai / DataDome / PerimeterX / Imperva?"
- Network advantage — Cloudflare sees ~20% of internet traffic. The ML model is trained on a data set no competitor has.
- Same edge as everything else — no separate inline agent, no traffic mirroring, no sidecar. If you're on Cloudflare, scoring is already running.
- TLS fingerprinting is structural — JA4 can't be spoofed without rewriting the client's TLS stack, and we use it as a scored signal plus cross-layer mismatch detection, not a brittle standalone block. One of the strongest signals in bot detection today.
- Verified bot list is cryptographic — no false-positives on Googlebot. Competitors often rely on User-Agent matching, which is trivial to spoof.
- One platform, one rule language — combine bot score with WAF, Rate Limiting, Access, API Shield. Not five vendors stitched together.
6. Close (~1 min)
"If you're already on Cloudflare, bot scoring is running right now — on every request hitting your zones. The only thing left is writing the rules that act on it."
Quick answers (if asked)
Will it block Googlebot?
No — verified bots are allowed automatically. Cryptographically verified, not header-based.
How accurate is the score, and how do the ranges work?
Trained on Cloudflare's global traffic. Cloudflare's own groupings: a score of 1 is Automated, 2–29 is Likely Automated, and 30–99 is Likely Human. A common enforcement pattern is Block or Managed Challenge below ~20 and let the rest through, but you set the line where it fits your risk.
Can attackers fake JA4?
Not easily. Unlike the old JA3, JA4 sorts ciphers and extensions before hashing, so you can't beat it by shuffling cipher order. Changing it means swapping the client's actual TLS library (curl-impersonate, a custom stack) — and if one parameter is off, the inconsistency itself flags them. Rotating it every request backfires too, since a fingerprint that keeps shifting is anomalous in its own right.
Should I just block by JA4?
Usually no. Millions of real users share a legit browser's JA4, so a blanket block causes mass false positives, and sophisticated bots rotate fingerprints anyway. Lead with a Bot Score rule (JA4 is already an input). Use JA4 in compound rules — block when JA4 says non-browser but the UA claims Chrome. Reserve a raw JA4 block for temporary, targeted mitigation during an active attack, then remove it.
Then why does the dashboard show the JA4 at all?
For explainability (audit why a request scored low), for Signals Intelligence (network-wide stats like browser_ratio per fingerprint), and for building precise compound rules — not for one-click blocking.
What about headless Chrome / Puppeteer?
Distinct fingerprints. Headless Chrome has missing browser APIs and a different TLS fingerprint than regular Chrome. We catch all major headless frameworks.
Does this work for APIs?
Yes — bot score works on any HTTP request. API Shield adds API-specific protections (schema validation, sequence checks, mTLS) on top.
How much?
Bot Fight Mode: free. Super Bot Fight Mode: included on Pro/Business. Full Bot Management: Enterprise add-on.
How long to deploy?
If already on Cloudflare — minutes. Scoring is already running, you just write rules. From scratch, a few hours to a day.
Will it slow down my site?
No. Scoring runs in the same pipeline as the WAF — single-digit milliseconds. No detour, no extra hops.
What's the false positive rate?
Industry-low. The model is conservative — it'd rather let a borderline request through than block a real customer. Most false positives are legitimate automation (uptime monitors, your own scripts) — allowlist by IP or fingerprint.
How is this different from a CAPTCHA?
Bot Management is invisible — most users never see anything. CAPTCHAs interrupt every user. When we do challenge, we use Turnstile (invisible CAPTCHA replacement), not picking out fire hydrants.
How do I tune it without breaking anything?
Start in log-only mode. Watch the score distribution for a week. Identify your own automation (monitoring, partners). Allowlist them. Then move to enforcement.
Does it work with Workers?
Yes — Workers can read cf.bot_management.score and react. Useful for custom block pages, dynamic responses, API logic.
What about logged-in users?
Bot score still applies, but session signals are factored in. A user logged in and active for 20 minutes is treated differently from a fresh low-score request to /login.
How do I block AI scrapers (OpenAI, Anthropic, Perplexity)?
AI Crawl Control — separate from regular Bot Management. Cloudflare maintains a list of identified AI crawlers. Allow, Block, Challenge, or Charge per AI bot. Available on all plans including Free.