> Source: https://txtfetch.com/extract/epub > Plain-text twin — every page on txtfetch.com has one. https://txtfetch.com/text --- # Ebooks in, chapter text out. EPUB's zipped-XHTML internals are exactly the kind of format archaeology txtfetch exists to hide from you. the-problem An EPUB is a zip archive of XHTML files, a manifest, and a spine that defines reading order. None of that is something you want to parse by hand just to get a book's text out. Generic text extractors expect a single flat document. They don't know what to do with a container format like this, and most teams don't have an EPUB parser sitting around. one-request-solution POST the .epub to txtfetch, and Apache Tika walks the manifest and spine for you. It returns the book's text in reading order as plain text, with the same request shape and response shape as every other format. curl ```curl curl -X POST https://api.txtfetch.com/v1/extract \ -H "Authorization: Bearer $TXTFETCH_KEY" \ -F file=@field-guide.epub ``` Python ```python import os import requests with open("field-guide.epub", "rb") as f: r = requests.post( "https://api.txtfetch.com/v1/extract", headers={"Authorization": f"Bearer {os.environ['TXTFETCH_KEY']}"}, files={"file": f}, ) print(r.json()["extracted_text"]) ``` JavaScript ```javascript import { readFile } from "node:fs/promises"; const file = new Blob([await readFile("field-guide.epub")]); const form = new FormData(); form.append("file", file, "field-guide.epub"); const res = await fetch("https://api.txtfetch.com/v1/extract", { method: "POST", headers: { Authorization: `Bearer ${process.env.TXTFETCH_KEY}` }, body: form, }); const { extracted_text } = await res.json(); console.log(extracted_text); ``` Go ```go package main import ( "bytes" "encoding/json" "fmt" "io" "mime/multipart" "net/http" "os" ) type extractResponse struct { Status string `json:"status"` ExtractedText string `json:"extracted_text"` } func main() { f, err := os.Open("field-guide.epub") if err != nil { panic(err) } defer f.Close() var body bytes.Buffer writer := multipart.NewWriter(&body) part, err := writer.CreateFormFile("file", "field-guide.epub") if err != nil { panic(err) } if _, err := io.Copy(part, f); err != nil { panic(err) } writer.Close() req, err := http.NewRequest("POST", "https://api.txtfetch.com/v1/extract", &body) if err != nil { panic(err) } req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY")) req.Header.Set("Content-Type", writer.FormDataContentType()) resp, err := http.DefaultClient.Do(req) if err != nil { panic(err) } defer resp.Body.Close() var result extractResponse if err := json.NewDecoder(resp.Body).Decode(&result); err != nil { panic(err) } fmt.Println(result.ExtractedText) } ``` Or skip the download. Pass a `url` parameter and txtfetch fetches the document server-side: curl ```curl curl -X POST "https://api.txtfetch.com/v1/extract?url=https://example.com/library/field-guide.epub" \ -H "Authorization: Bearer $TXTFETCH_KEY" ``` Python ```python import os import requests r = requests.post( "https://api.txtfetch.com/v1/extract", headers={"Authorization": f"Bearer {os.environ['TXTFETCH_KEY']}"}, params={"url": "https://example.com/library/field-guide.epub"}, ) print(r.json()["extracted_text"]) ``` JavaScript ```javascript const endpoint = new URL("https://api.txtfetch.com/v1/extract"); endpoint.searchParams.set("url", "https://example.com/library/field-guide.epub"); const res = await fetch(endpoint, { method: "POST", headers: { Authorization: `Bearer ${process.env.TXTFETCH_KEY}` }, }); const { extracted_text } = await res.json(); console.log(extracted_text); ``` Go ```go package main import ( "encoding/json" "fmt" "net/http" "net/url" "os" ) type extractResponse struct { Status string `json:"status"` ExtractedText string `json:"extracted_text"` } func main() { endpoint, err := url.Parse("https://api.txtfetch.com/v1/extract") if err != nil { panic(err) } q := endpoint.Query() q.Set("url", "https://example.com/library/field-guide.epub") endpoint.RawQuery = q.Encode() req, err := http.NewRequest("POST", endpoint.String(), nil) if err != nil { panic(err) } req.Header.Set("Authorization", "Bearer "+os.Getenv("TXTFETCH_KEY")) resp, err := http.DefaultClient.Do(req) if err != nil { panic(err) } defer resp.Body.Close() var result extractResponse if err := json.NewDecoder(resp.Body).Decode(&result); err != nil { panic(err) } fmt.Println(result.ExtractedText) } ``` ``` { "status": "success", "extracted_text": "..." } ``` what-comes-back The corpus behind [/diff](https://txtfetch.com/diff) has no recorded .epub document yet, so we have nothing honest to show you here. Run one of your own instead. The [free converter](https://txtfetch.com/tools/file-to-text) reads the file in your browser, and nothing is uploaded. formats-covered - `.epub` faq **How do I extract text from an EPUB file?**: POST the .epub as multipart form data to https://api.txtfetch.com/v1/extract with your API key in the Authorization header. You get back { "status": "success", "extracted_text": "..." }, with the book's text in reading order. **What about older Kindle formats like .mobi or .azw?**: txtfetch's core support targets EPUB. .mobi and .azw run through the same endpoint and often extract cleanly via Tika's fallback parsers, but EPUB is the best-tested path today. **Does it preserve chapter order?**: Yes. Tika reads the EPUB's spine (its defined reading order) rather than the arbitrary order files happen to be zipped in. go-further - [Drop a real .epub and see the extracted text, chapter by chapter, free →](https://txtfetch.com/tools/epub-to-text) - [API quickstart →](https://txtfetch.com/docs) - [How the pipeline works →](https://txtfetch.com/how-it-works) - [Extract it from your language →](https://txtfetch.com/for) - [Get an API key →](https://app.txtfetch.com/signup) Books & archives - [`.zip`](https://txtfetch.com/extract/zip) - [All formats →](https://txtfetch.com/extract) ## Send a real .epub through it. One HTTP call returns the text. Read one in your browser first, for free. [Get an API key →](https://app.txtfetch.com/signup) [Open the free .epub reader →](https://txtfetch.com/tools/epub-to-text)