InvarixBlogBenchmark
Benchmark12 min read

Invarix.Guard vs DeBERTa-v3 and Microsoft Presidio: a reproducible head to head benchmark

Invarix.Guard matches the open source DeBERTa-v3 injection model at comparable latency and beats Presidio on PII F1 with zero false positives on that test set. In process, no Python sidecar.

On this page · 01/08

We benchmarked Invarix.Guard against the strongest open source baseline for each task:

  • DeBERTa-v3 prompt injection model. ProtectAI/deberta-v3-base-prompt-injection-v2, served from a FastAPI sidecar. It is also the prompt injection backbone inside Invarix.Guard, so both sides of this comparison run the same weights.
  • Microsoft Presidio. The de facto open source PII analyzer, spaCy NER backbone.

3 benchmarks across 2,450 multilingual prompts and 210 synthetic PII rows. All datasets are public, and the methodology is described in full below.

Summary

Table 1Headline verdict per benchmark · Invarix.Guard v1.2.0
BenchmarkVerdict
Multilingual prompt injection (2,450 rows, 16 langs vs DeBERTa-v3)F1 0.769 vs 0.778 (-0.9pp). FPR tied. P50 comparable (189ms vs 226ms)
Adversarial robustness (TextBugger / WordDeletion / CaseFlip)Both detectors retain 94-100% TPR.
PII detection (210 synthetic rows vs Microsoft Presidio)F1 0.928 vs 0.897 (+3.1pp). FPR 0% vs 12%

Beats Microsoft Presidio on PII detection F1 with zero false positives on this test set. Matches the open source DeBERTa-v3 prompt injection model on detection metrics at comparable latency. Survives adversarial perturbation as well as the dedicated comparator. All of it runs in process in .NET, with no Python sidecar.

Why this benchmark

We wanted to know whether Invarix.Guard, running injection and PII detection in process in a .NET app, holds up against the dedicated tools.

This isn't a benchmark against a closed commercial product we can't reach. It's a benchmark against the strongest open source baselines for each task, run on the same input we'd run through Invarix.Guard. If you're evaluating commercial guards, run them on the same fixture: the harness, per row results, and rerun commands are public.

Methodology

Scoring.Binary. A guard "flags" an input when it returns Block or Flag (Invarix), is_valid: false (the DeBERTa sidecar), or any entity above 0.5 confidence (Presidio). TPR = flagged / actual positives. FPR = flagged / actual negatives.

1. Multilingual prompt injection

Datasets (all public). 200 English attacks from HumanCompatibleAI/tensor-trust-data + 900 multilingual attacks from walledai/MultiJail (ICLR 2024) covering Arabic, Bengali, Italian, Javanese, Korean, Swahili, Thai, Vietnamese, Chinese (100 each) + 300 NLLB-200 translated TensorTrust extraction attacks in German, Spanish, French, Japanese, Portuguese, Russian (50 each) + 150 PromptBench style perturbed extraction attacks (see §2). Clean prompts: 200 English from databricks/databricks-dolly-15k + 700 multilingual from OpenAssistant/oasst1. Total: 2,450 rows, 16 languages.

Comparator: ProtectAI/deberta-v3-base-prompt-injection-v2 loaded via HuggingFace transformers in a small FastAPI sidecar.

Table 2Aggregate metrics · 2,450 row prompt injection benchmark
GuardTPRFPRF1AccP50P95
Invarix.Guard v1.2.067.8%14.9%0.76974.2%189ms589ms
DeBERTa-v3 (FastAPI sidecar)69.2%15.0%0.77875.0%226ms396ms

Detection metrics are deterministic and reproduce byte identical across runs. Latency is system dependent. Across two runs on the same fixture, Invarix's P50 ranged 129-189ms, DeBERTa's 226-256ms. Read the latency numbers as comparable, not as a win.

Aggregate metrics across the 2,450 row prompt injection benchmark
Fig. 1Aggregate metrics across the full 2,450 row benchmark.

Per language attack catch rate

Per language attack catch rate (TPR), Invarix vs DeBERTa
Fig. 2Per language attack catch rate. Invarix vs DeBERTa.

Where both detectors are strong: the NLLB translated extraction set (de/es/fr/ja/pt/ru) lands at 94-100% TPR for both guards. Translation preserves enough injection signal that detection holds up. Bengali (98/100) is the only professionally translated MultiJail language where both detectors stay strong.

Where both detectors fall below 50%: Italian (18/15), Chinese (4/5), Korean (22/26), Vietnamese (35/38), Swahili (41/41), and Javanese (48/43). Arabic (67/75) and Thai (87/95) are middling. Both pipelines wrap the same model, so this is that model's ceiling on MultiJail, not an Invarix limitation.

Per language false positive rate

Per language false positive rate, Invarix vs DeBERTa
Fig. 3Per language clean false positive rate. Lower is better.

OASST1 is volunteer collected and not fully curated. Some of the "clean" Spanish / French / Portuguese rows look like unlabeled jailbreak attempts: both guards flag the same rows, which points to real attacks rather than independent false positives. The headline numbers don't try to correct for that.

2. Adversarial robustness

We applied 3 PromptBench style perturbations to 50 English TensorTrust extraction attacks each, generating 150 adversarial rows:

  • TextBugger. Random character insertions / deletions / swaps + Cyrillic homoglyph substitutions (10% of letter chars).
  • WordDeletion. Drops 20% of non stopword tokens.
  • CaseFlip. Randomly upper/lowercases 30% of letters.
Table 3Detection rate per attack family · 400 original rows, 50 per perturbation
Attack familynInvarix TPRDeBERTa TPR
extraction (original)40097%98%
extraction-CaseFlip50100%100%
extraction-TextBugger5094%94%
extraction-WordDeletion5094%94%
Detection rate per attack family under PromptBench style perturbations
Fig. 4Detection rate per attack family under PromptBench style perturbations.

Both detectors are perturbation robust. Only a 3-4pp drop from the original 97-98% catch rate even under character edits + Cyrillic homoglyph injection + 20% word deletion. Under CaseFlip both stayed at 100% (regex case insensitive + ML normalizes case anyway).

3. PII detection (vs Microsoft Presidio)

Fixture:111 synthetic PII rows + 99 clean rows = 210 total. PII types covered: EMAIL, PHONE, SSN, CREDIT_CARD, IP, NAME, ADDRESS, DATE_OF_BIRTH. All "real looking" PII values are fabricated: RFC 2606 example.com domains, 555 area code phones, 192.0.2.0/24 IPs, IIN-9999 Luhn valid credit cards, 9XX area SSNs. Fixture seed = 42 for reproducible generation.

Comparator: Microsoft Presidio (presidio-analyzer 2.2.362), spaCy en_core_web_lg backbone, via local FastAPI sidecar.

Table 4Aggregate metrics · 210 row PII benchmark
GuardTPRFPRF1AccP50
Invarix.Guard v1.2.086.5%0.0%0.92892.9%109ms
Microsoft Presidio90.1%12.1%0.89789.0%10ms
Aggregate metrics for the PII benchmark
Fig. 5Aggregate metrics, 210 row PII benchmark.

Invarix wins on F1 (0.928 vs 0.897) and accuracy (92.9% vs 89.0%), with 0% false positives against Presidio's 12%. Presidio wins on latency by an order of magnitude (~11× faster). It is a single task NER pipeline, while Invarix runs the full guard stack (PII + injection + toxicity + safety) on every input. Disable the scanners you don't need via GuardOptions.EnablePiiDetection / EnableContentSafety / EnablePromptInjection and latency drops roughly linearly with the number of scanners you leave enabled.

Per PII type detection rate

Per PII type detection rate, Invarix vs Presidio
Fig. 6Per PII type detection rate. Invarix vs Presidio.
Table 5Per PII type detection rate · row contains type counted
PII typeInvarixPresidioΔ
EMAIL100%100%—
IP100%100%—
SSN100%100%—
ADDRESS100%100%—
CREDIT_CARD100%58%+42pp
PHONE100%80%+20pp
NAME83%97%-14pp
DATE_OF_BIRTH0%100%-100pp

The DATE_OF_BIRTH gap is by design in Invarix. Bare dates require a "born / dob / birthday" context word to flag, avoiding false positives on log lines, calendar entries, and historical dates. Presidio flags any date shaped string. So last seen 03/15/2024 raises no PII alarm in Invarix, but Presidio flags it. The 12% FPR difference is largely from this.

A note on latency

Invarix.Guard's P50 (189ms on the prompt injection benchmark, 109ms on the PII benchmark) is the cost of running four ML scannerson every input: prompt injection (DeBERTa-v3), PII NER (XLM-RoBERTa multilingual), embedding based content safety (e5-base), and toxicity (multilingual DistilBERT). Each comparator runs one scanner instead of four. The DeBERTa sidecar (226ms P50) is slowed by the IPC hop and PyTorch's eager mode inference. Presidio (10ms P50) is faster because spaCy NER is leaner.

We tried parallelizing the four ML scanners via Parallel.Invoke(independent ONNX sessions, each thread safe per Microsoft.ML.OnnxRuntime docs). Across several runs the change made no measurable difference. Threadpool dispatch overhead matched the parallelism win when individual scanners were already fast (clean inputs, ~30-100ms each). The slow tail is concentrated on long adversarial inputs where SentencePiece tokenization runs O(n) regardless of inference, so reordering doesn't help.

The architectural fixes that would actually move P95 are bigger projects: a multi task model (one inference for all scanners), smaller or quantized backbones, and token level input capping. They are on the roadmap, not in v1.2.

Caveats and known limitations

  • Translated attack data. The EU language attacks are machine translated. NLLB-200-distilled-600M preserves injection intent but may not reflect the natural distribution of native speaker attacks.
  • Attack coverage.Multilingual coverage is uneven. MultiJail's languages don't overlap with OASST1's clean prompt languages. Only Chinese appears in both sets, so it is the one language with a per language F1.
  • Clean set contamination. OASST1 contains some unlabeled jailbreak attempts in non English splits, inflating the FPR for both guards on those languages.
  • Sample size. n=50-100 per language is small. Confidence intervals at this size are ±5pp. Per language differences within ±3pp should be read as ties.
  • Synthetic fixture. The PII fixture is built from hand crafted templates with deterministic fakes. Real world PII patterns may differ.
  • Shared model. The same DeBERTa model is wrapped in both prompt injection pipelines. This is a benchmark of the integration around the model, not of which model is better.
  • System dependent latency. Same fixture, same machine, two runs on different days produced P50 differences of ~30% on the prompt injection benchmark. Detection metrics, by contrast, reproduced byte identical. Treat the latency numbers above as one data point, not a precise measurement.

Run it yourself

Every detection number above reproduces from the committed harness: fixtures, per row result JSON, and rerun commands are in the public repo under benchmarks/. If you want the same stack in your own process, the Community tier is free: dotnet add package Invarix.Guard. It runs in process in your ASP.NET Core app, and nothing you scan leaves the box.