> Source: https://txtfetch.com/for/go > Plain-text twin — every page on txtfetch.com has one. https://txtfetch.com/text --- # The library Go's ecosystem doesn't have. Python and Node both have an imperfect but usable Office parser. Go doesn't. Its OCR story means cgo, and cgo breaks the static, cross-compiled binary Go programs are supposed to be. the-parser-zoo Go's document-parsing ecosystem is thin compared to Python's or Node's. The gap is structural, not just a missing package. PDF text extraction has options. ledongthuc/pdf covers basic digital-native text. UniDoc/unipdf is capable, but it is commercially licensed for anything beyond noncommercial use. Neither is a python-docx or mammoth equivalent with meaningful adoption for reading .docx, .pptx, or .xlsx as plain text. Teams either shell out to a converter, or hand-parse the OOXML zip-of-XML themselves. Searching for a golang docx parser turns up low-adoption repos and wrappers around non-Go tooling, not a python-docx equivalent. OCR is the sharper problem. The practical option is gosseract, a cgo binding to the system libtesseract. That means CGO\_ENABLED=1, a matching Tesseract install on every build and deploy target, and no more single static binary. It also means giving up the easy cross-compilation Go is usually chosen for. A Lambda or scratch-container Go binary that needs OCR faces a choice. Either bundle a Tesseract shared library manually, or give up the zero-dependency deploy story entirely. | library | covers | stops at | | --- | --- | --- | | `ledongthuc/pdf` | Digital-native PDF text extraction | No OCR, no Office formats; struggles with complex layouts | | `unipdf (UniDoc)` | PDF text, forms, and more, capably | Commercial license required beyond noncommercial/eval use | | `gosseract` | OCR via cgo bindings to libtesseract | Requires CGO\_ENABLED=1 and a matching native Tesseract install on every build/deploy target | | `(community OOXML readers)` | Fragmented, low-adoption .docx/.xlsx readers | No mammoth/python-docx-level standard; most teams hand-parse the zip themselves | one-request txtfetch replaces the whole gap with one request. POST a file, or pass ?url=, to https://api.txtfetch.com/v1/extract using net/http from the standard library. There is no cgo, no CGO\_ENABLED=1, and no Tesseract shared library to vendor alongside your binary. The response is the same { "status": "success", "extracted\_text": "..." }, whether the source was a PDF, an Office file, or a scan. A Go service stays a single static binary, regardless of what it's asked to extract. Go ```go package main import ( "bytes" "encoding/json" "fmt" "io" "mime/multipart" "net/http" "os" ) type extractResponse struct { Status string `json:"status"` ExtractedText string `json:"extracted_text"` } func main() { f, err := os.Open("report.pdf") if err != nil { panic(err) } defer f.Close() var body bytes.Buffer writer := multipart.NewWriter(&body) part, err := writer.CreateFormFile("file", "report.pdf") if err != nil { panic(err) } if _, err := io.Copy(part, f); err != nil { panic(err) } writer.Close() req, err := http.NewRequest("POST", "https://api.txtfetch.com/v1/extract", &body) if err != nil { panic(err) } req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY")) req.Header.Set("Content-Type", writer.FormDataContentType()) resp, err := http.DefaultClient.Do(req) if err != nil { panic(err) } defer resp.Body.Close() var result extractResponse if err := json.NewDecoder(resp.Body).Decode(&result); err != nil { panic(err) } fmt.Println(result.ExtractedText) } ``` ``` { "status": "success", "extracted_text": "..." } ``` from-a-url Skip the download entirely. Pass a `url` parameter and txtfetch fetches the document server-side: Go ```go package main import ( "encoding/json" "fmt" "net/http" "net/url" "os" ) type extractResponse struct { Status string `json:"status"` ExtractedText string `json:"extracted_text"` } func main() { endpoint, err := url.Parse("https://api.txtfetch.com/v1/extract") if err != nil { panic(err) } q := endpoint.Query() q.Set("url", "https://example.com/report.pdf") endpoint.RawQuery = q.Encode() req, err := http.NewRequest("POST", endpoint.String(), nil) if err != nil { panic(err) } req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY")) resp, err := http.DefaultClient.Do(req) if err != nil { panic(err) } defer resp.Body.Close() var result extractResponse if err := json.NewDecoder(resp.Body).Decode(&result); err != nil { panic(err) } fmt.Println(result.ExtractedText) } ``` errors Every non-success response carries a stable `error.code`. Match on that, not on `error.message`. See the full [error reference](https://txtfetch.com/docs/errors) for every code and HTTP status txtfetch can return. Go ```go package main import ( "bytes" "encoding/json" "fmt" "io" "mime/multipart" "net/http" "os" ) type extractError struct { Code string `json:"code"` Message string `json:"message"` } type extractResponse struct { Status string `json:"status"` ExtractedText string `json:"extracted_text"` Error *extractError `json:"error"` } func main() { f, err := os.Open("report.pdf") if err != nil { panic(err) } defer f.Close() var body bytes.Buffer writer := multipart.NewWriter(&body) part, err := writer.CreateFormFile("file", "report.pdf") if err != nil { panic(err) } if _, err := io.Copy(part, f); err != nil { panic(err) } writer.Close() req, err := http.NewRequest("POST", "https://api.txtfetch.com/v1/extract", &body) if err != nil { panic(err) } req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY")) req.Header.Set("Content-Type", writer.FormDataContentType()) resp, err := http.DefaultClient.Do(req) if err != nil { panic(err) } defer resp.Body.Close() var result extractResponse if err := json.NewDecoder(resp.Body).Decode(&result); err != nil { panic(err) } switch { case resp.StatusCode == 200: fmt.Println(result.ExtractedText) case resp.StatusCode == 429: fmt.Printf("back off: %s, retry after %ss\n", result.Error.Code, resp.Header.Get("Retry-After")) default: fmt.Printf("extraction failed: %s — %s\n", result.Error.Code, result.Error.Message) } } ``` big-files-and-batches Large uploads or slow documents are routed to an async job automatically. That returns a `202` plus a `job_id` to poll, and `?async=true` forces that path for any request. Direct upload size ceilings by plan: Hobby 10 MB, Developer 50 MB, Scale 200 MB. Use `?url=` for anything larger. Server-side fetches aren't held to the upload ceiling. See [async jobs & webhooks](https://txtfetch.com/docs/async) for the full lifecycle, including webhook delivery instead of polling. Go ```go package main import ( "encoding/json" "fmt" "net/http" "net/url" "os" "time" ) type jobResponse struct { Status string `json:"status"` JobID string `json:"job_id"` ExtractedText string `json:"extracted_text"` } func poll(req *http.Request) (*jobResponse, error) { resp, err := http.DefaultClient.Do(req) if err != nil { return nil, err } defer resp.Body.Close() var result jobResponse if err := json.NewDecoder(resp.Body).Decode(&result); err != nil { return nil, err } return &result, nil } func main() { token := "Bearer " + os.Getenv("TXTFETCH_KEY") endpoint, _ := url.Parse("https://api.txtfetch.com/v1/extract") q := endpoint.Query() q.Set("url", "https://example.com/report.pdf") q.Set("async", "true") endpoint.RawQuery = q.Encode() submitReq, _ := http.NewRequest("POST", endpoint.String(), nil) submitReq.Header.Set("Authorization", token) job, err := poll(submitReq) if err != nil { panic(err) } pollURL := "https://api.txtfetch.com/v1/extract/" + job.JobID var result *jobResponse for { pollReq, _ := http.NewRequest("GET", pollURL, nil) pollReq.Header.Set("Authorization", token) result, err = poll(pollReq) if err != nil { panic(err) } if result.Status != "processing" { break } time.Sleep(2 * time.Second) } fmt.Println(result.ExtractedText) } ``` gotchas - CGO\_ENABLED=0 builds are the common way to produce a truly static, scratch-container-friendly Go binary. gosseract's cgo dependency means those builds can't use it at all. - Cross-compiling a cgo-dependent binary (say, building for linux/arm64 from an amd64 CI runner) means cross-compiling libtesseract too, not just your Go code. That is a build-pipeline problem most Go teams would rather not own. - There's no single OOXML reader with the adoption level of Python's python-docx. Most Go codebases either shell out to LibreOffice or Pandoc for conversion, or hand-walk the .docx zip's document.xml. - unipdf's licensing (commercial past noncommercial/evaluation use) is worth checking before it ends up load-bearing in a shipped product. formats - [PDF](https://txtfetch.com/extract/pdf) - [Word / PowerPoint / Excel](https://txtfetch.com/extract/docx) - [Spreadsheets](https://txtfetch.com/extract/xlsx) - [Scans & images (OCR)](https://txtfetch.com/extract/image) faq **How do I extract text from a PDF in Go without cgo?**: POST the PDF (or pass ?url=) to https://api.txtfetch.com/v1/extract using net/http. There is no cgo, no CGO_ENABLED flag, and no native library to link. The response is { "status": "success", "extracted_text": "..." }, whether the PDF is digital-native or scanned. **Is there a golang library for reading .docx or .xlsx files?**: Nothing has the adoption level of Python's python-docx or openpyxl. Most Go teams shell out to a converter, or hand-parse the OOXML zip. txtfetch handles the whole Office family (.docx, .pptx, .xlsx, and the legacy .doc, .xls, .ppt) through the same endpoint as everything else. **How do I OCR a scanned document in Go without gosseract's cgo dependency?**: POST the scan to txtfetch. OCR runs server-side, so your Go binary never links libtesseract and never sets CGO_ENABLED=1. It stays cross-compilable to any target, the same way it was before you needed OCR. **Does txtfetch have an official Go SDK?**: No, and it doesn't need one. It's a single JSON-in, JSON-out HTTP request. The standard library's net/http and encoding/json cover the whole client, in well under fifty lines, shown above. go-further - [How to extract text from a PDF for RAG →](https://txtfetch.com/blog/extract-text-from-pdf-for-rag) - [Document workflows & automation →](https://txtfetch.com/solutions/document-workflows) - [Measured extraction accuracy by category →](https://txtfetch.com/benchmarks) - [API quickstart →](https://txtfetch.com/docs) - [How the pipeline works →](https://txtfetch.com/how-it-works) - [Get an API key →](https://app.txtfetch.com/signup) other-languages - [`Python`](https://txtfetch.com/for/python) - [`JavaScript`](https://txtfetch.com/for/javascript) - [`Java`](https://txtfetch.com/for/java) - [`C# / .NET`](https://txtfetch.com/for/csharp) - [All languages →](https://txtfetch.com/for) ## Paste it into your project. The Go snippet above runs as written. Add your key and it works. [Get an API key →](https://app.txtfetch.com/signup) [More SDK quickstarts →](https://txtfetch.com/docs/quickstarts)