Extract text from a document
Extract text on the Canvas Perch server, or choose an external service when a scanned PDF needs OCR.
Follow the steps
Upload a supported document into Library and open its item page.
Choose Extract text locally to process the document on the Canvas Perch server without an external OCR service.
If a scanned PDF needs OCR, choose Scanned PDF options, select an available service, and review the disclosure.
Choose Allow this document only to send the entire PDF to that service for this job. Usage charges may apply.
Derive text and colors locally
A narrated walkthrough with English captions and a full transcript.
Read the transcript
For an image, choose Extract palette. Canvas Perch derives the colors from the image itself, so the same pixels produce the same result.
Select more than one image to extract a combined palette for the collection.
For a supported document, choose Extract text locally. The job stays visible while it runs, and the text is saved with the item when processing finishes.
A scanned PDF can offer hosted OCR separately, only after you approve that document.
What to know
- External OCR requires your consent for each PDF, an available service you select, and a file up to 8 MB.
- A hosted result records its provider and is shown with a reminder to compare the transcription with the original document.
- The app checks visible hosted text and returned reasoning fields for Han characters before saving. This is a conservative character rule that also catches Japanese Kanji and Korean Hanja; it is not a complete language detector.
- Read Privacy and processing for external-service and retention details.
Troubleshooting
- If local extraction fails without an OCR-needed result, fix the source file or retry after checking the job error; hosted fallback is not an automatic retry.
- If an OCR service is unavailable, choose another available service or Keep it local. Smart tagging is separate: an eligible new capture may send a limited amount of extracted text to an external AI service. Turn smart tags off in Settings → Notifications before extraction if you do not want that text sent for tagging.
- If the hosted result is rejected for language policy or exceeds the output limit, no partial text is saved and the job reports the failure.
- If the file is too large, split the PDF before requesting hosted OCR. The normal upload limit and the hosted PDF limit are separate.

