Originally published at or the uses a two-step flow. First, you submit a PDF URL and receive a check ID. Second, you retrieve the result using that ID. Both steps use standard HTTP with Bearer token authentication.
The document must be publicly accessible — either a direct URL to the file, or a presigned URL from your storage provider (S3, Google Cloud Storage, Cloudflare R2, Vercel Blob). HTPBE? downloads the PDF server-side, so the URL only needs to be temporarily accessible.
cURL
CODE# Step 1 — submit the PDF URL for analysis
curl -X POST https://api.htpbe.tech/v1/analyze \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://your-bucket.s3.amazonaws.com/documents/contract.pdf"}'
Response:
CODE{"id": "ck_9f4a2e1b-3d7c-4a8e-b1f2-9e0d3c5a7b8f"}
CODE# Step 2 — retrieve the verdict
curl https://api.htpbe.tech/v1/result/ck_9f4a2e1b-3d7c-4a8e-b1f2-9e0d3c5a7b8f \
-H "Authorization: Bearer YOUR_API_KEY"
Response (abbreviated):
CODE{
"id": "ck_9f4a2e1b-3d7c-4a8e-b1f2-9e0d3c5a7b8f",
"status": "modified",
"modification_confidence": "high",
"modification_markers": ["INCREMENTAL_UPDATES", "PRODUCER_MISMATCH"],
"xref_count": 4,
"has_digital_signature": false,
"creator": "Microsoft Word",
"producer": "iLovePDF",
"creation_date": 1704067200,
"modification_date": 1709251200
}
The
modification_markersarray tells you exactly which signals triggered the verdict — not just that the document is suspect, but why.
Python
CODEimport os
import httpx # pip install httpx
API_KEY = os.environ["HTPBE_API_KEY"]
BASE_URL = "https://api.htpbe.tech/v1"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
def verify_pdf(pdf_url: str) -> dict:
"""Submit a PDF URL and return the full analysis result."""
# Step 1: submit
submit_response = httpx.post(
f"{BASE_URL}/analyze",
headers=HEADERS,
json={"url": pdf_url},
timeout=30,
)
submit_response.raise_for_status()
check_id = submit_response.json()["id"]
# Step 2: retrieve result
result_response = httpx.get(
f"{BASE_URL}/result/{check_id}",
headers=HEADERS,
timeout=30,
)
result_response.raise_for_status()
return result_response.json()
def route_document(pdf_url: str) -> str:
"""Return an action based on the PDF forensic verdict."""
result = verify_pdf(pdf_url)
status = result["status"]
if status == "intact":
return "accept"
elif status == "modified":
markers = result.get("modification_markers", [])
print(f"Tampering detected: {', '.join(markers)}")
return "reject"
else: # inconclusive
# Consumer software origin — route to manual review
return "manual_review"
The
httpxlibrary is used here because it has a cleaner API thanrequestsfor JSON workflows, butrequestsworks identically — replacehttpx.postwithrequests.postandhttpx.getwithrequests.get.
Handling errors in Python
CODEimport httpx
def verify_pdf_safe(pdf_url: str) -> dict | None:
try:
return verify_pdf(pdf_url)
except httpx.HTTPStatusError as e:
status = e.response.status_code
if status == 401:
raise RuntimeError("Invalid HTPBE API key") from e
if status == 402:
raise RuntimeError("HTPBE subscription required") from e
if status == 422:
# URL did not return a valid PDF
return None
raise
except httpx.TimeoutException:
# Handle timeout — retry or queue for later
return None
Node.js / TypeScript
CODEconst API_KEY = process.env.HTPBE_API_KEY!;
const BASE_URL = 'https://api.htpbe.tech/v1';
interface HTPBEResult {
id: string;
status: 'intact' | 'modified' | 'inconclusive';
modification_confidence: 'certain' | 'high' | 'none' | null;
modification_markers: string[];
xref_count: number;
has_digital_signature: boolean;
modifications_after_signature: boolean;
signature_removed: boolean;
creator: string | null;
producer: string | null;
creation_date: number | null;
modification_date: number | null;
}
async function verifyPdf(pdfUrl: string): Promise<HTPBEResult> {
// Step 1: submit
const submitRes = await fetch(`${BASE_URL}/analyze`, {
method: 'POST',
headers: {
Authorization: `Bearer ${API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ url: pdfUrl }),
});
if (!submitRes.ok) {
const body = await submitRes.json().catch(() => ({}));
throw new Error(`HTPBE submit failed ${submitRes.status}: ${JSON.stringify(body)}`);
}
const { id } = await submitRes.json() as { id: string };
// Step 2: retrieve result
const resultRes = await fetch(`${BASE_URL}/result/${id}`, {
headers: { Authorization: `Bearer ${API_KEY}` },
});
if (!resultRes.ok) {
throw new Error(`HTPBE result fetch failed ${resultRes.status}`);
}
return resultRes.json() as Promise<HTPBEResult>;
}
// Production usage example
async function handleDocumentSubmission(pdfUrl: string): Promise<'accept' | 'reject' | 'review'> {
const result = await verifyPdf(pdfUrl);
switch (result.status) {
case 'intact':
return 'accept';
case 'modified':
console.log('Tampering markers:', result.modification_markers);
if (result.modifications_after_signature) {
console.log('Document was modified after digital signing');
}
return 'reject';
case 'inconclusive':
// Consumer software origin — may be legitimate, route to human
return 'review';
}
}
Reading the result: what each field means
The fields that matter most for routing decisions:
Field
Type
What it tells you
status
intact/modified/inconclusive
The primary verdict
modification_confidence
certain/high/none
How confident the verdict is
modification_markers
string[]
Which specific signals triggered the verdict
modifications_after_signature
boolean
Content added after a valid digital signature
signature_removed
boolean
A digital signature was stripped from the document
xref_count
number
Number of edit sessions in the file
creator
string
Software that created the document
producer
string
Software that last processed the document
The three
certainconfidence markers — meaning the confidence level is"certain"rather than"high":
MODIFICATIONS_AFTER_SIGNATURE— cryptographically verifiable
SIGNATURE_REMOVED— the signature slot exists but the signature is gone
DIFFERENT_DATES— creation and modification dates differ by more than 15 seconds in an impossible sequence
Everything else produces
"high"confidence, not"certain". For workflows where false positives are costly (legal proceedings, for example),certainmarkers warrant automatic rejection whilehighmarkers might warrant manual review.
Routing theinconclusiveverdict
inconclusivedoes not mean the document is untrustworthy. It means the document was created with consumer software — Microsoft Word, Google Docs, LibreOffice, Canva — and lacks the structural patterns of institutionally-generated documents.
The right routing depends on what you are processing:
Documents that claim institutional origin (bank statements, tax certificates, court filings, insurance policies):
inconclusiveshould be treated the same asmodified. A bank statement claiming to come from a financial institution should not be produced by Google Docs.
User-generated documents (forms, applications, letters, CVs):
inconclusiveis expected and acceptable. Route to normal processing.
CODEfunction shouldReject(result: HTPBEResult, claimsInstitutionalOrigin: boolean): boolean {
if (result.status === 'modified') return true;
if (result.status === 'inconclusive' && claimsInstitutionalOrigin) return true;
return false;
}
Testing without real documents
All HTPBE? plans include a test API key. Test keys accept mock URLs that return predictable responses — similar to Stripe test cards. Use these in your test suite to cover every verdict branch without consuming production quota.
CODE# Clean document — returns status: intact
https://api.htpbe.tech/v1/test/clean.pdf
# Tampered document — returns status: modified
https://api.htpbe.tech/v1/test/modified-high.pdf
# Consumer software origin — returns status: inconclusive
https://api.htpbe.tech/v1/test/inconclusive.pdf
# Signature removed — returns status: modified, signature_removed: true
https://api.htpbe.tech/v1/test/signature-removed.pdf
# Modified after signing — returns modifications_after_signature: true
https://api.htpbe.tech/v1/test/modified-medium.pdf
Keep your test key in a
.env.testfile and never let it touch production flows.
What this does not detect
Two scenarios where forensic metadata analysis has limits:
Documents created fraudulently from scratch. If someone fabricates a bank statement using the same software a real bank uses, generates plausible timestamps, and produces a structurally consistent PDF — the file may pass analysis. Forensic analysis catches editing of existing documents and lazy fabrication. Sophisticated forgery from scratch, using professional tools, may require additional signals (visual content analysis, issuer fraud detection).
Encrypted documents. Strongly encrypted PDFs cannot be analyzed for structural signals. The analysis will flag this as inconclusive by necessity.
For most operational workflows — invoice processing, loan applications, recruitment document checks — forensic metadata analysis catches the overwhelming majority of attempted fraud, which uses off-the-shelf PDF editors rather than sophisticated fabrication tools.
↗ Original-Artikel auf dev.to lesenVollständiger Original-BerichtAusführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
Ähnliche Beiträge
Auch interessante Nachrichten Detect PDF Tampering Programmatically: Developer Guide
Thematisch verwandte Begriffe: Detect, Tampering, Programmatically, Developer · 6 Treffer
Claude Code Observability with OpenTelemetry
Hackers Exploit PaperCut NG/MF Flaws to Steal Credentials and Deploy Meterpreter
Magento and Adobe Commerce StyleSmuggler 0-Day RCE Actively Exploited in Attacks
BPFDoor Scanner
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
SOCIAL SHARE CARD GENERATOR