https://txtfetch.com/for/go/
The library Go's ecosystem doesn't have.
Python and Node both have an imperfect but usable Office parser. Go doesn't. Its OCR story means cgo, and cgo breaks the static, cross-compiled binary Go programs are supposed to be.
the-parser-zoo
Go's document-parsing ecosystem is thin compared to Python's or Node's. The gap is structural, not just a missing package. PDF text extraction has options. ledongthuc/pdf covers basic digital-native text. UniDoc/unipdf is capable, but it is commercially licensed for anything beyond noncommercial use. Neither is a python-docx or mammoth equivalent with meaningful adoption for reading .docx, .pptx, or .xlsx as plain text. Teams either shell out to a converter, or hand-parse the OOXML zip-of-XML themselves. Searching for a golang docx parser turns up low-adoption repos and wrappers around non-Go tooling, not a python-docx equivalent. OCR is the sharper problem. The practical option is gosseract, a cgo binding to the system libtesseract. That means CGO_ENABLED=1, a matching Tesseract install on every build and deploy target, and no more single static binary. It also means giving up the easy cross-compilation Go is usually chosen for. A Lambda or scratch-container Go binary that needs OCR faces a choice. Either bundle a Tesseract shared library manually, or give up the zero-dependency deploy story entirely.
| library | covers | stops at |
|---|---|---|
ledongthuc/pdf | Digital-native PDF text extraction | No OCR, no Office formats; struggles with complex layouts |
unipdf (UniDoc) | PDF text, forms, and more, capably | Commercial license required beyond noncommercial/eval use |
gosseract | OCR via cgo bindings to libtesseract | Requires CGO_ENABLED=1 and a matching native Tesseract install on every build/deploy target |
(community OOXML readers) | Fragmented, low-adoption .docx/.xlsx readers | No mammoth/python-docx-level standard; most teams hand-parse the zip themselves |
one-request
txtfetch replaces the whole gap with one request. POST a file, or pass ?url=, to https://api.txtfetch.com/v1/extract using net/http from the standard library. There is no cgo, no CGO_ENABLED=1, and no Tesseract shared library to vendor alongside your binary. The response is the same { "status": "success", "extracted_text": "..." }, whether the source was a PDF, an Office file, or a scan. A Go service stays a single static binary, regardless of what it's asked to extract.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
)
type extractResponse struct {
Status string `json:"status"`
ExtractedText string `json:"extracted_text"`
}
func main() {
f, err := os.Open("report.pdf")
if err != nil {
panic(err)
}
defer f.Close()
var body bytes.Buffer
writer := multipart.NewWriter(&body)
part, err := writer.CreateFormFile("file", "report.pdf")
if err != nil {
panic(err)
}
if _, err := io.Copy(part, f); err != nil {
panic(err)
}
writer.Close()
req, err := http.NewRequest("POST", "https://api.txtfetch.com/v1/extract", &body)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY"))
req.Header.Set("Content-Type", writer.FormDataContentType())
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
var result extractResponse
if err := json.NewDecoder(resp.Body).Decode(&result); err != nil {
panic(err)
}
fmt.Println(result.ExtractedText)
}{
"status": "success",
"extracted_text": "..."
}from-a-url
Skip the download entirely. Pass a url parameter and txtfetch fetches the document server-side:
package main
import (
"encoding/json"
"fmt"
"net/http"
"net/url"
"os"
)
type extractResponse struct {
Status string `json:"status"`
ExtractedText string `json:"extracted_text"`
}
func main() {
endpoint, err := url.Parse("https://api.txtfetch.com/v1/extract")
if err != nil {
panic(err)
}
q := endpoint.Query()
q.Set("url", "https://example.com/report.pdf")
endpoint.RawQuery = q.Encode()
req, err := http.NewRequest("POST", endpoint.String(), nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY"))
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
var result extractResponse
if err := json.NewDecoder(resp.Body).Decode(&result); err != nil {
panic(err)
}
fmt.Println(result.ExtractedText)
}errors
Every non-success response carries a stable error.code. Match on that, not on error.message. See the full error reference for every code and HTTP status txtfetch can return.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
)
type extractError struct {
Code string `json:"code"`
Message string `json:"message"`
}
type extractResponse struct {
Status string `json:"status"`
ExtractedText string `json:"extracted_text"`
Error *extractError `json:"error"`
}
func main() {
f, err := os.Open("report.pdf")
if err != nil {
panic(err)
}
defer f.Close()
var body bytes.Buffer
writer := multipart.NewWriter(&body)
part, err := writer.CreateFormFile("file", "report.pdf")
if err != nil {
panic(err)
}
if _, err := io.Copy(part, f); err != nil {
panic(err)
}
writer.Close()
req, err := http.NewRequest("POST", "https://api.txtfetch.com/v1/extract", &body)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY"))
req.Header.Set("Content-Type", writer.FormDataContentType())
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
var result extractResponse
if err := json.NewDecoder(resp.Body).Decode(&result); err != nil {
panic(err)
}
switch {
case resp.StatusCode == 200:
fmt.Println(result.ExtractedText)
case resp.StatusCode == 429:
fmt.Printf("back off: %s, retry after %ss\n", result.Error.Code, resp.Header.Get("Retry-After"))
default:
fmt.Printf("extraction failed: %s — %s\n", result.Error.Code, result.Error.Message)
}
}big-files-and-batches
Large uploads or slow documents are routed to an async job automatically. That returns a 202 plus a job_id to poll, and ?async=true forces that path for any request. Direct upload size ceilings by plan: Hobby 10 MB, Developer 50 MB, Scale 200 MB. Use ?url= for anything larger. Server-side fetches aren't held to the upload ceiling. See async jobs & webhooks for the full lifecycle, including webhook delivery instead of polling.
package main
import (
"encoding/json"
"fmt"
"net/http"
"net/url"
"os"
"time"
)
type jobResponse struct {
Status string `json:"status"`
JobID string `json:"job_id"`
ExtractedText string `json:"extracted_text"`
}
func poll(req *http.Request) (*jobResponse, error) {
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
var result jobResponse
if err := json.NewDecoder(resp.Body).Decode(&result); err != nil {
return nil, err
}
return &result, nil
}
func main() {
token := "Bearer " + os.Getenv("TXTFETCH_KEY")
endpoint, _ := url.Parse("https://api.txtfetch.com/v1/extract")
q := endpoint.Query()
q.Set("url", "https://example.com/report.pdf")
q.Set("async", "true")
endpoint.RawQuery = q.Encode()
submitReq, _ := http.NewRequest("POST", endpoint.String(), nil)
submitReq.Header.Set("Authorization", token)
job, err := poll(submitReq)
if err != nil {
panic(err)
}
pollURL := "https://api.txtfetch.com/v1/extract/" + job.JobID
var result *jobResponse
for {
pollReq, _ := http.NewRequest("GET", pollURL, nil)
pollReq.Header.Set("Authorization", token)
result, err = poll(pollReq)
if err != nil {
panic(err)
}
if result.Status != "processing" {
break
}
time.Sleep(2 * time.Second)
}
fmt.Println(result.ExtractedText)
}gotchas
- CGO_ENABLED=0 builds are the common way to produce a truly static, scratch-container-friendly Go binary. gosseract's cgo dependency means those builds can't use it at all.
- Cross-compiling a cgo-dependent binary (say, building for linux/arm64 from an amd64 CI runner) means cross-compiling libtesseract too, not just your Go code. That is a build-pipeline problem most Go teams would rather not own.
- There's no single OOXML reader with the adoption level of Python's python-docx. Most Go codebases either shell out to LibreOffice or Pandoc for conversion, or hand-walk the .docx zip's document.xml.
- unipdf's licensing (commercial past noncommercial/evaluation use) is worth checking before it ends up load-bearing in a shipped product.
formats
faq
- How do I extract text from a PDF in Go without cgo?
- POST the PDF (or pass ?url=) to https://api.txtfetch.com/v1/extract using net/http. There is no cgo, no CGO_ENABLED flag, and no native library to link. The response is { "status": "success", "extracted_text": "..." }, whether the PDF is digital-native or scanned.
- Is there a golang library for reading .docx or .xlsx files?
- Nothing has the adoption level of Python's python-docx or openpyxl. Most Go teams shell out to a converter, or hand-parse the OOXML zip. txtfetch handles the whole Office family (.docx, .pptx, .xlsx, and the legacy .doc, .xls, .ppt) through the same endpoint as everything else.
- How do I OCR a scanned document in Go without gosseract's cgo dependency?
- POST the scan to txtfetch. OCR runs server-side, so your Go binary never links libtesseract and never sets CGO_ENABLED=1. It stays cross-compilable to any target, the same way it was before you needed OCR.
- Does txtfetch have an official Go SDK?
- No, and it doesn't need one. It's a single JSON-in, JSON-out HTTP request. The standard library's net/http and encoding/json cover the whole client, in well under fifty lines, shown above.
go-further
other-languages
Paste it into your project.
The Go snippet above runs as written. Add your key and it works.
Get an API key →