There is a class of projects that teaches you more about a language than any tutorial ever could. Building a PDF engine from scratch in Go is one of them. It is not glamorous. It is not trendy. But it forces you to confront memory management, binary serialization, concurrency safety, interface design, and performance profiling all at once, in a domain where correctness is non-negotiable.
This article walks through the lessons learned building GoPdfSuit (~500 Github ⭐), a production PDF engine written in Go that generates 1.5 million financial PDFs in roughly 45 minutes on a single node, achieves PDF/A-4 and PDF/UA-2 compliance, and exposes itself as a REST API, a Go library, and Python CGO bindings simultaneously.
Note: While I have six years of overall experience including two years working specifically with Go, I rarely encountered these types of challenges in my day-to-day work, as my role focused primarily on implementing new features within an existing architecture. Working on gopdfsuit was an excellent learning experience; it allowed me to dive deep into performance optimization and taught me a great deal. Additionally, I utilized AI tools to assist in the development of this project. Below are some of the key takeaways.
Building GoPdfSuit from a blank editor to a production-grade PDF engine—shipping PDF 2.0, PDF/A-4, PDF/UA-2, PKCS#7 signing, and secure redaction—forced a shift from standard business logic to hardcore systems engineering. To achieve ~2,000+ aggregate ops/s on a mixed financial workload (48 workers, PDF/A enabled) with sub-10ms generation, we bypassed framework debates to fight the Go allocator, optimize cache lines, and strictly implement ISO 32000 semantics. To extend this performance globally, I exported the core engine as gopdflib, building a high-performance Python wrapper via CGO to deliver native-speed PDF operations to Python environments.
1. Optimization
The Art of Micro-Optimizations
Working on GoPdfSuit forced a shift from "business logic" to "systems engineering." The kinds of optimizations that are rarely needed in day-to-day feature work += on []byte with sprintf equivalents, direct byte append operations instead of fmt.Sprintf, and bit-shift approximations for division became the norm here.
Key techniques used across the codebase:
appendTextForPDFZero-alloc text encoding directly into[]byte, eliminatingstringintermediates on everyTjPDF operator. Found ininternal/pdf/utils.go, used acrossdraw.gocell rendering paths.
Byte-scratch buffers Hot paths use stack-fixed[24]byteand[12]bytescratch buffers for numeric formatting.appendFmtNumavoidsstrconv.AppendFloat, documented as ~10% CPU savings in profiling.
RuneSet bitmap Replacedmap[rune]boolfor tracking used characters with a dense 64 KiB bitmap (font/runeset.go), cutting map inserts on the font subsetting hot path.
Fast alpha blending Replaced integer division (/ 255) per pixel component with fast multiplication:((r*a + white) * 0x8081) >> 23. Found ininternal/pdf/image.go.
256-byte lookup tables for hex encoding instead offmt.Sprintf, and boundary check elimination via pre-sized buffers.
Batched writes Reduced ~25K separateWritecalls for a 5K-cell table down to ~5K by batching PDF commands per cell. Implemented ininternal/pdf/draw.go.
Pre-grow all buffers Page content streams pre-grow to 64 KiB, compress buffers to 64 KiB, assembly buffer to 64 KiB avoiding incrementalappendgrowth during hot generation.
The Four-Phase Performance Program
The optimization journey was structured across 4 passes with 41 total tasks, documented across guides/cursor/PR_PERFORMANCE_OPTIMIZATION.md, guides/cursor/PASS1_BLUEPRINTS.md through PASS4_OPTIMIZATION_PLAN.md, and guides/additionalnotes/PERFORMANCE_OPTIMIZATIONS.md:
| Pass | Focus | Tasks | Key Outcomes |
|---|---|---|---|
| Pass 1 | Low-hanging fruit | 10 | Buffer pooling, zero-alloc text encoding, batched writes, RuneSet bitmap, image cache singleflight |
| Pass 2 | Architecture | 12 | Direct-write APIs, parallel decode/compress, incremental MD5, sparse CIDToGIDMap, macro benchmarks |
| Pass 3 | Advanced memory | 5 | Allocation-free WrapTextInto, typed []StructKid replacing []interface{}, redact parser unification |
| Pass 4 | Load-test hotspots | 14 | PDF/UA gating, final PDF slice pool, StructElem pool, template pool, parallel zlib, p99 fixes |
Headline Results:
2061 ops/s peak (1705 ops/s 10-run average) on Zerodha gold-standard workload (48 workers, PDF/A + tagged + signatures) ~197% faster than Go 1.24 baseline, ~4.4× faster than Go 1.26
- Serial 2000-row PDF/A: ~31–36 ms/op, ~163K allocs/op (~46% fewer than pre-optimization)
- HTTP load test: ~5.7× throughput, ~18× faster p99, −88% heap in-use (442 MB → 55 MB)
memclrCPU under load: 49.7% → 27.0% (−46% relative)
PDF File Size Optimization
Several techniques are employed to keep PDF output small without sacrificing generation speed:
FlateDecode compression All content streams, font streams, ICC profiles, and metadata are compressed using zlib.zlib.NewWritercarries a ~256 KB compression table cost, so a centralsync.Poolof zlib writers is shared across generation paths infont/compression.go.
Compression level balancingzlib.BestSpeedis used to balance size vs. performance. The compression pool pre-grows buffers to 64 KiB (max(4096, len/4)before zlib).
Font subsetting A complete TrueType/OpenType subsetting engine (876 lines ininternal/pdf/font/subset.go) extracts only the glyphs actually used in the document. It handles composite glyph dependencies (recursively pulling in component glyphs), remaps glyph IDs, rebuilds all required TTF tables (head,hhea,maxp,glyf,loca,hmtx,cmap,post,name), and generates sparse CIDToGIDMap entries. The font registry tracks character usage viaMarkCharsUsed()across all drawing paths (including signatures), then callsGenerateSubsets()at finalize time.
Image deduplication An FNV-1a hash-basedimageCachewithsync.RWMutexskips redundant PNG decoding and compression when the same image appears multiple times in a document. Backed bysingleflightto deduplicate concurrent decodes of the same image hash.
Concurrency & Memory Pools
Seven active sync.Pool instances across the codebase:
| Pool | Size | Purpose |
|---|---|---|
pdfBufferPool | 64 KB pre-grow | PDF assembly *bytes.Buffer |
finalPDFSlicePool | 256 KB cap | Scratch []byte for final PDF assembly |
scratchBufPool | 128 B | Small strconv scratch buffers |
rgbDataPool | 1 MB | RGB image conversion buffers |
structElemPool | PDF/UA structure tree *StructElem nodes | |
templatePDFPool | HTTP handler *models.PDFTemplate instances | |
ZlibWriterPool / CompressBufPool | 64 KB | Zlib compression writers and buffers |
2. PDF: Not Just a Normal File
PDF as a Complex Format (Similar to HTML but bit a complex)
A PDF is far more complex than a simple file format it's closer to a programmatic document description language that shares conceptual similarities with HTML but operates at a much lower level. While HTML describes structure and relies on browser engines for layout, a PDF must define every glyph position, color space, font embedding, and encryption detail in binary form.
PDF Structure Deep Dive
The internal structure of a PDF 2.0 document (ISO 32000-2) as generated by GoPdfSuit:
Header%PDF-2.0magic bytes with optional binary comment marker
Body Sequence of indirect objects (numbered 1..N), each with a generation number, dictionary, and optional stream:
Catalog (root object): References pages tree, outlines, structure tree, metadata, output intents, viewer preferences
Pages Tree: Hierarchical page nodes with/Kids,/Count,/MediaBox
Page Objects:/Contentsstream (PDF operators),/Resources(fonts, XObjects),/Annots(links, signatures)
Font Objects: Type0 fonts with CIDFontType2 + Identity-H encoding, ToUnicode CMaps, font descriptor with metrics, font file streams
XObject Images: DCTDecode/FlateDecode streams with/ColorSpace,/BitsPerComponent,/Width,/Height
Structure Tree (PDF/UA):/StructTreeRoot→/Karray of struct elements with/S(type: Document, Table, TR, TD, P, H1-H6...),/Pg(page reference),/K(MCID references)
Outline Tree:/Outlines→ bookmark hierarchy with/Title,/Dest,/Count
Metadata: XMP stream under/Metadatawith Dublin Core, PDF/A, PDF/UA extension schemas
Output Intents:/OutputIntentsarray with ICC profile stream and/DestOutputProfile
Signature: PKCS#7 CMS signature with/ByteRange,/Contents(placeholder + signature value)
Cross-Reference Table Compact xref subsections mapping object numbers to byte offsets, with/W,/Index, and compressed object streams (ObjStm) for non-contiguous ranges
Trailer/Size,/Root,/Info,/IDarray,/Prevfor incremental updates,startxrefpointer
Key implementation details:
- Objects use fixed ID allocation starting from 2000+: catalog → pages → streams → fonts
- Content streams emit standard PDF operators:
BT/ET(text blocks),Tj/TJ(text),Tm(matrix),Tf(font),BDC/EMC(marked content for structure tree),q/Q(graphics state save/restore),re/f(rectangles),Do(XObjects),cm(coordinate transforms) - The xref writer is shared across generation, merge, and XFDF paths in
internal/pdf/xref/xref.go
- Redaction reads existing PDFs by scanning
N G obj … endobjpatterns, expanding ObjStm streams, and augmenting with xref-stream parsing
PDF Image Encoding
Image handling in internal/pdf/image.go (609 lines):
Supported formats: PNG (decoded viaimage/png), JPEG (passthrough viaimage/jpeg), SVG (parsed and rendered viainternal/pdf/svg/svg.go)
Color space handling: Images are converted to DeviceRGB for PDF/A, with an explicit embedding of a custom-built sRGB ICC v2.1 profile (buildSRGBICCProfileinpdfa.go) with hand-corrected TRC curves to prevent washed-out output in Adobe Acrobat
Compression: DCTDecode for JPEG (with color transform), FlateDecode for RGB rasters
Alpha blending: Fast pixel-level alpha compositing using((r*a + white) * 0x8081) >> 23(the 0x8081 magic number approximates division by 255)
Deduplication: FNV-1a hash-based image cache avoids re-decoding and re-compressing duplicate images
Coordinate transforms: SVG import applies explicit flip matrix (1/w 0 0 -1/h 0 1 cm) to map SVG coordinates to PDF bottom-left user space
PDF Coloring
Color spaces supported: DeviceRGB, DeviceGray, ICCBased RGB, ICCBased Gray
PDF/A requirement: All DeviceRGB/Gray streams reference an ICC profile via/OutputIntentsthe engine hand-builds valid ICC v2.1 profiles from scratch (sRGB with corrected TRC curves, Gray with proper whitepoint)
CMYK: Not implemented financial templates are RGB-first
Cell coloring: Table cells support background colors applied viaq/Qpairs aroundre(rectangle) andf(fill) operators
PDF Layout
Layout uses a top-down Y model internally while emitting standard PDF bottom-left user space:
PageManager.CurrentYPos = height - topMargin(ininternal/pdf/pagemanager.go)- Table rendering in
draw.go(~1800+ lines) handles column widths, text wrapping, row heights, cell borders, superscripts/subscripts, checkboxes, placeholders, and auto-column-width detection - Text wrapping uses
WrapTextIntowith runninglineWidthtracking and reusable[][]byteline buffers to avoid per-line allocations - Line width calculations use real TTF
hmtx/glyfmetrics for custom fonts and hard-coded Standard 14 width tables for WinAnsi fonts
PDF Metadata & Headers for Compliance
XMP metadata (internal/pdf/metadata.go,internal/pdf/pdfa.go): Generates PDF/A-4 + PDF/UA-2 compliant XMP with Dublin Core (dc:title,dc:creator,dc:description), XMP Media Management (xmp:CreateDate,xmp:ModifyDate), PDF/A extension schema (pdfaid:part=4,pdfaid:rev=2020), PDF/UA extension schema (pdfuaid:part=2,pdfuaid:rev=2024)
Catalog entries:/MarkInfo << /Marked true >>(required for PDF/UA),/Lang (en-US),/ViewerPreferences << /DisplayDocTitle true >>,/StructTreeRoot,/OutputIntents
Document ID: Two-part/IDarray with random byte-generated file IDs in the trailer
Trailer compliance: PDF 2.0 trailers include/ID,/Info(optional in 2.0),/Size,/Root
3. Compliance Hard to Start, Easier with Implementation
Compliance seemed intimidating until the implementation was underway. Understanding the specifications deeply and building the infrastructure piece by piece made it progressively easier.
PDF/A-4 Compliance (ISO 19005-4:2020)
GoPdfSuit targets PDF/A-4 (part=4, rev=2020), the latest archiving standard based on PDF 2.0:
All fonts must be embedded Custom fonts are fully embedded and subsetted. Standard fonts (Helvetica, Courier, Times) are substituted with Liberation font equivalents (LiberationSans,LiberationMono,LiberationSerif) when PDF/A mode is enabled, handled ininternal/pdf/font/pdfa.go. The Liberation fonts are downloaded on demand (with double-check caching to avoid races), then embedded and subsetted.
ICC color profiles required A valid sRGB ICC v2.1 profile is constructed from scratch (buildSRGBICCProfileinpdfa.go, 507 lines). The TRC curves are manually corrected to avoid the washed-out rendering that stock profiles produce in Adobe Acrobat.
XMP metadata mandatory Generated with all required extension schemas (PDF/A ID, PDF/UA ID, Dublin Core).
No encryption in PDF/A (pdfaCompliant+Security.Enabled= rejection)
No external references All resources are embedded.
PDF/UA-2 Compliance (ISO 14289-2:2024)
PDF/UA-2 (part=2, rev=2024) requires fully tagged PDFs with accessibility structure:
Structure tree (internal/pdf/structure.go, 436 lines): Implements 25+ standard structure types: Document, Part, Sect, Div, H1-H6, P, L, LI, Lbl, LBody, Table, TR, TH, TD, Figure, Caption, Form, Link, Reference.
Marked content: Every content element is wrapped inBDC/EMCoperators with MCID (Marked Content ID) references. TheStructureManagertracks MCID allocation, manages the ParentTree, and creates link/annotation structure elements.
Tagged PDF opt-in: TheTaggedPDFconfig flag gates all structure tree construction when off, a no-opStructureManageravoids allocation overhead entirely (Pass 4 optimization).
Bookmark sect elements:CreateBookmarkSect()generates section structure elements for PDF/UA-2 structure destinations.
Arlington Compliance with XML Structure Trees
The ArlingtonCompatible config flag enables PDF 2.0 Arlington Model compliance this ensures full font metrics are available, the structure tree conforms to the Arlington TSV specification, and all required dictionary entries are present. The Arlington model provides a machine-readable specification of all valid PDF object keys and value types.
veraPDF Validator
The verapdf/ directory contains the veraPDF validation tool for automated compliance checking. Community feedback from real users pointed to veraPDF as the gold standard for validating PDF/A and PDF/UA compliance.
4. Engagement with Real Users
Running an open-source project meant interacting directly with users and incorporating their feedback:
User feedback loop: Real users tested the library in production scenarios and provided actionable feedback on features like XFDF form filling, redaction behavior, and PDF/A validation.
Tooling suggestions: Community members recommended veraPDF as the validator for checking PDF/A and PDF/UA compliance leading to theverapdf/integration in the repository.
Issue-driven development: Features like PDF splitting, secure redaction with OCR integration, HTML-to-PDF conversion, and the Python CGO bindings were driven by user requests.
Documentation feedback: The React playground, API documentation, and benchmark comparisons were shaped by what users actually needed to understand to adopt the library.
5. pprof & Performance Profiling
Profiling Infrastructure
pprof is deeply integrated into the project for both development and production diagnostics:
Server-side endpoints (internal/handlers/handlers.go):
/debug/pprof/Index
/debug/pprof/cmdlineCommand line args
/debug/pprof/profile30-second CPU profile
/debug/pprof/symbolSymbol table
/debug/pprof/traceExecution trace
/debug/pprof/heapHeap profile
/debug/pprof/goroutineGoroutine dump
/debug/pprof/allocsAllocation profile
/debug/pprof/blockBlocking profile
/debug/pprof/mutexMutex contention profile
/debug/pprof/threadcreateThread creation profile
All endpoints are gated to localhost only (127.0.0.1, ::1) for security.
Opt-in heap dump on exit (cmd/gopdfsuit/main.go): ENABLE_PROFILING=1 writes a heap profile to /tmp/mem.prof on server shutdown.
Benchmark pprof (sampledata/benchmarks/gopdflib/run_pprof_bench.sh):
- 5000 iterations, 48 workers
- 1 timing run + 5 CPU profile runs + 1 heap profile run
- Profiles saved as
.proffiles underguides/cursor/baselines/gopdflib_pprof_runs/
CPU Hotspot Analysis (from pprof results)
Top CPU hotspots identified and addressed across passes:
| Hotspot | Initial | After Optimization | Fix |
|---|---|---|---|
drawTable (cumulative) | ~37% | ~17.73% | Hoisted scratch buffers, batched writes, P4-04 |
memclrNoHeapPointers (flat) | 49.7% (under load) | 27.0% (under load) | Buffer pre-grow, pooling, P4-03 |
compress/flate | ~20% | ~5-8% | Zlib writer pool, P1-04, P4-08 |
image/png.readImagePass | 21.7% | Varies | Image cache + singleflight, P1-10 |
| PNG decoding | Hot | Eliminated for dupes | FNV-1a hash cache, P1-10 |
BeginMarkedContentBuf | ~6.8% | Reduced | Tagged PDF gating (P4-01) when not needed |
Heap Hotspot Analysis
From 5000-iteration pprof benchmark:
bytes.growSlice: 443.40 MB (59% of total) addressed by pre-growing buffers
compress/flate.NewWriter: 88.34 MB cumulative addressed by zlib pooling
GenerateTemplatePDF: 642.64 MB cumulative largest single consumer- Under HTTP load: heap in-use reduced from 442 MB → 55 MB (−88%)
6. Architecture
Design Patterns in Practice
Several design patterns emerged naturally from the PDF engine's requirements:
1. Facade Pattern pkg/gopdflib/ provides a clean public API surface that delegates to internal implementation packages. All public types are type aliases (type PDFTemplate = models.PDFTemplate), keeping the surface minimal:
GeneratePDF(template)→pdf.GenerateTemplatePDF(template)
MergePDFs(files)→merge.MergePDFs(files)
SplitPDF(file, spec)→merge.SplitPDF(file, spec)
FillPDFWithXFDF(pdf, xfdf)→form.FillPDFWithXFDF(pdf, xfdf)
ConvertHTMLToPDF(req)→ HTML-to-PDF via Chrome headless
2. Builder Pattern OutlineBuilder (internal/pdf/outline.go, 505 lines) constructs the PDF outline tree with a fluent API: NewOutlineBuilder(pm, encryptor) → BuildOutlines(bookmarks).
3. Factory Pattern NewPDFAHandler(config, pageManager, encryptor) and signature.NewPDFSigner(config) encapsulate complex object construction with dependencies.
4. Strategy / Adapter Patterns via Interfaces:
ObjectEncryptorinterface allows switching between AES-128, AES-256, RC4, and no-op encryption without changing callers
SignaturePageContextinterface decouples the signature subsystem fromPageManagerinternals
OCRProviderinterface allows plugging in different OCR backends for redaction
signatureContextAdapteradapts*PageManagerto implementSignaturePageContextwithout creating circular dependencies
5. Object Pool Pattern (sync.Pool) Seven pools (detailed in Section 1) heavily reduce GC pressure on hot paths.
6. Registry / Singleton Pattern CustomFontRegistry with GetFontRegistry() provides system-wide font management. Thread safety is achieved through CloneForGeneration() each PDF generation gets a shallow clone with isolated usage tracking and noLock: true to avoid mutex overhead on single-threaded generation paths.
7. Component Pattern PDF structure is built from typed element components (Table, Spacer, Image, Footer, Title, Bookmark) assembled via ordered Element slices in the template.
Decoupled Architecture for a PDF Engine
Key architectural decisions:
Data flows one way: Template → Parser → PageManager → ContentStreams → Assembly → Final PDF
Font registry is cloned per generation eliminates mutex contention, makes concurrent generation safe
Parallelism is gated behindruntime.NumCPU()semaphore middleware incmd/gopdfsuit/main.goprevents goroutine thrashing
Per-page zlib compression is parallel (errgroup) but assembly, encryption, and xref writing stay serial for deterministic object numbering
context.Contextis not used in the PDF pipeline no cancellation, no value chains; keeps the hot path lightweight
7. One Project, Multiple Technologies
Go Backend (Gin Web App)
The project serves a full web application via the Gin framework:
Entry point:cmd/gopdfsuit/main.go
Framework: Gin (release mode) with custom lightweight panic recovery (avoidsgin.Recovery()'s per-request defer overhead)
Concurrency control: Semaphore middleware sized toruntime.NumCPU()prevents goroutine thrashing (originally 100 goroutines on 24 cores caused massive context-switch overhead)
Routes: Serves the Vite-built React SPA fromdocs/, plus 14 API endpoints under/api/v1for PDF generation, merging, splitting, XFDF filling, HTML-to-PDF, redaction, font management, and OCR
pprof: Full profiling exposed on localhost-only/debug/pprof/endpoints
Middleware: CORS (allowing GitHub Pages origin), Google OAuth (Cloud Run only), semaphore-based concurrency gating
Go PDF Library (gopdflib)
The public Go library at pkg/gopdflib/ exposes all PDF operations as a clean API:
GeneratePDF(template)Main generation entry point
MergePDFs(files)Merge multiple PDFs into one
SplitPDF(file, spec)Split PDF by page ranges
FillPDFWithXFDF(pdfBytes, xfdfBytes)Fill form fields
ConvertHTMLToPDF(req)HTML to PDF via headless Chrome
ConvertHTMLToImage(req)HTML to raster image
GetAvailableFonts()List registered fonts
GetFontRegistry()Access font system for advanced use- Redaction API
ExtractTextPositions,FindTextOccurrences,ApplyRedactions,ApplyRedactionsAdvanced
React Frontend (Vite SPA)
A full single-page application built with React 18+ and Vite:
12 pages: Home (landing), Editor (template builder), Viewer, Merge, Split, Filler (XFDF), HtmlToPdf, HtmlToImage, Comparison (benchmarks), Documentation, Redaction, Screenshots
Componentization: Each page is a self-contained route-level component with reusable sub-components (Navbar,PdfPreview,PerformanceSection,AuthGuard,Toast,BackgroundAnimation, editor components)
Routing: React Router v6 withHashRouterfor GitHub Pages compatibility
Auth integration:AuthGuardcomponent gates the Editor route behind Google OAuth on Cloud Run deployments
UI framework: MUI (Material UI) components with a custom theme
Build output: Vite builds intodocs/which is served by the Go backend as static assets
8. CGO Python Bindings (gopdflib → pypdfsuit)
The entire Go PDF engine is exported as a Python package via CGO shared library:
Implementation (bindings/python/cgo/exports.go, 437 lines):
- Compiles to a C shared library (
.so/.dylib) usinggo build -buildmode=c-shared
- Exports 14 C-callable functions using
//exportdirectives andimport "C":
GeneratePDF,MergePDFs,SplitPDF,FillPDFWithXFDF
ConvertHTMLToPDF,ConvertHTMLToImage
GetAvailableFonts,GetPageInfo,ExtractTextPositions
FindTextOccurrences,ApplyRedactions,ApplyRedactionsAdvanced
ParsePageSpec
- Memory management:
FreeBytesResultandFreeBytesArrayResultfor caller-side cleanup - Python package:
bindings/python/pypdfsuit/withsetup.py/pyproject.tomlfor PyPI distribution - Build scripts:
build.sh(Linux/macOS) andbuild.bat(Windows) for compiling the shared library
This means the same high-performance Go engine powers both Go and Python ecosystems without any runtime penalty from inter-language communication overhead only the initial function call crosses the CGO boundary, and then all PDF generation happens natively in Go memory.
9. Beating the Industry Standard with a FOSS Project
Zerodha Benchmark Comparison
Zerodha (India's largest retail brokerage) publicly documented their infrastructure for generating 1.5 million digitally signed PDF contract notes daily using a 40-node Nomad cluster running Typst/LaTeX CLI tools achieving approximately 1,000 PDFs/sec aggregate throughput.
GoPdfSuit achieves comparable or superior performance:
| Metric | Zerodha (40 nodes) | GoPdfSuit (1 node) | Improvement |
|---|---|---|---|
| Throughput (peak) | ~1,000 ops/s | 2,061 ops/s | 2× on single node |
| Throughput (10-run avg) | ~1,000 ops/s | 1,705 ops/s | 1.7× on single node |
| Per-core efficiency | ~1.6 PDFs/sec/core | ~86 PDFs/sec/core | ~54× more efficient |
| Infrastructure | 40 nodes | 1 node (24 vCPUs) | 40× fewer nodes |
| Time for 1.5M PDFs | 25 min (40 nodes) | ~15 min (1 node) | Faster on 1/40th hardware |
Cost Analysis
| Architecture | Required Nodes | Hourly Cost (AWS) | Monthly Cost | Savings |
|---|---|---|---|---|
| Zerodha (Typst) | ~40 instances | ~$24.50/hr | ~$306.00 | |
| gopdflib | 2 instances | ~$1.84/hr | ~$23.00 | ~92% |
Why is gopdflib so fast?
Native binary generation Generates PDF binary structure directly in RAM, no external process spawning.
Zero IO overhead No temporary files, no disk I/O; streams bytes directly in memory.
Goroutine concurrency Thousands of lightweight goroutines saturate all cores without OS thread overhead.
Asset reuse Font subsets and image assets are processed once and reused across millions of documents.
10. Maintaining a Live Open-Source Project
Building and maintaining GoPdfSuit as a public open-source project (~500 GitHub stars) involved:
Repository structure: Clean separation between library (pkg/gopdflib/), engine (internal/pdf/), web app, benchmarks, guides, and bindings
CI documentation: Guides for deployment, release checklists, troubleshooting, and Python porting
Community interaction: GitHub Issues and user feedback driving feature development
Version management: Go module with proper versioning (v5), PyPI package for Python bindings
Licensing: MIT license for both the Go library and Python bindings
Live playground: GitHub Pages deployment of the React playground atchinmay-sawant.github.io/gopdfsuit
15+ historical optimization logs inguides/14_02_Optimizations/documenting the iterative performance journey
Agent-assisted development: Cursor AI-generated pass blueprints (guides/cursor/PASS1_BLUEPRINTS.mdthroughPASS4_OPTIMIZATION_PLAN.md) layered on top of measured benchmarks not instead of them
11. GCloud Deployments
GCP Deployment & Architecture
Hands-on GCP Learning: Gained practical, end-to-end experience deploying a Go and React application on Google Cloud Platform.
Strategic Architecture Decisions: Conducted a thorough analysis and research phase comparing Google App Engine and Cloud Run to determine the optimal deployment strategy based on cost, scalability, and performance.
Resource Optimization: Settled on a dual-deployment approach using App Engine’s F1 instance class for standard hosting alongside tailored Cloud Run instances (512 MiB memory ceiling) optimized viaK_SERVICEenvironment detection.
Project Configuration
App Engine Standard Setup: Configured theapp.yamlarchitecture to manage runtime environments (go124), strict autoscaling limits, custom entry points, and essential environment variables (Google OAuth, Cloud Run URLs, and Vite configurations).
React Frontend Integration: Unified the React SPA (built via Vite) with the Go server binary, configuring Gin'sStaticFSmiddleware to serve static assets alongside custom SPA fallback routing.
Security & CORS: Implemented Google OAuth middleware to gate specific routes and configured precise CORS permissions to allow seamless communication between the GitHub Pages frontend and the Cloud Run API.
12. Building and Deploying Docker Image via Multi-Stage Docker Build
Multi-Stage Docker Image
The dockerfolder/Dockerfile uses a 2-stage build:
Docker Hub Publishing
The image is published to Docker Hub (linked from the makefile via gdocker-push target), making it available for anyone to pull and run.
Cloud Run Optimized Variant
Dockerfile_cloudrun mirrors the same 2-stage pattern but builds specifically for Cloud Run:
13. React Project Management
What Was Learned About React
The frontend started as a basic HTML page and evolved into a full React SPA, teaching:
Componentization:
- Each PDF operation (Editor, Merge, Split, Filler, HtmlToPdf, Redaction) is a route-level page component
- Reusable UI components:
Navbar(navigation with mobile-responsive hamburger),PdfPreview(embedded PDF viewer),PerformanceSection(benchmark visualization),AuthGuard(conditional OAuth wrapper),Toast(notification system),BackgroundAnimation(visual polish) - Editor sub-components: Form fields, template builder, JSON editor, live preview
- Home page sub-components: Hero section, Features grid, QuickStart guide, API overview, Comparison preview, Footer
Frontend Layout:
- Responsive design using MUI (Material UI) components with a custom theme
- Mobile-friendly navigation with collapsible menu
- Hash-based routing (
HashRouter) for GitHub Pages compatibility (no server-side URL rewriting needed) - SPA architecture with client-side routing for all 12 pages
React Hooks & Patterns:
useState,useEffectfor component state and side effects
useContextfor shared state (auth context, theme)
useReffor DOM element references (file inputs, preview containers)
useNavigatefor programmatic navigation- Custom hooks for API interaction and loading states
vite.config.jsfor build optimization and environment variable injection- API configuration via
utils/apiConfigwith environment-aware host selection
Project Management:
- Vite as the build tool (fast dev server, optimized production builds)
- Environment-based configuration (
.env.examplefor templating) - npm scripts for development, building, and previewing
- Build output integrated with the Go backend serving (output directory:
docs/)
What started as document rendering became systems engineering
GoPdfSuit is not a thin wrapper around a C library—it is a native generator (internal/pdf/generator.go orchestrates fonts, structure trees, encryption, and signing) with selective read/modify paths (merge, XFDF, redact).
What initially started as a simple XFDF parser side project quickly evolved into a fully compliant PDF engine supporting strict PDF/A-4 and PDF/UA-2 standards. The entire transformation took approximately 3 months of work—a hyper-fast timeline made possible by leaning heavily into AI pair programming (specifically GitHub Copilot, Antigravity, and more recently, Cursor and OpenCode).
By moving away from bloated, licensed enterprise solutions, this native Go engine doesn't just cut compute overhead—it represents a potential cost savings of $2,000 to $4,000 ARR per deployment, not just for my organization, but for any team swapping out commercial PDF infrastructure.
The lesson that repeated every pass: stop guessing, profile everything, respect the allocator, and treat ISO 32000 as a contract you test with real PDFs and VeraPDF-minded compliance tests—not with wishful string building.
Takeaways for Go Engineers
If you are building high-performance Go services, steal the discipline: pooled zlib, typed structure trees, honest benchmarks that state worker count, and docs that separate peak from average from serial micro-bench.
Check it out & Benchmark It 🚀
We’re right on the cusp of 500 stars on GitHub! ⭐ If you find the architecture or the performance numbers useful, drop a star to help us cross the finish line and keep the momentum going.
Run the Zerodha and internal/pdf benchmarks yourself, and share what you measure on your hardware:
Repository: github.com/chinmay-sawant/gopdfsuit
Live docs & playground: chinmay-sawant.github.io/gopdfsuit