Skip to content

Request IDs and retries ​

The requestId ​

Each extraction takes a requestId of yours: up to 128 characters, with letters, digits, _, -, . and :. Use one per document, with no personal data in it; a new UUID for each certificate works well.

Repeating without paying twice ​

The same requestId, with the same file and options, within 15 minutes of the answer, gets the same answer at no charge and without running the engines, with the header Idempotent-Replayed: true. With another file or other options, the same requestId makes a new request.

While the first request still runs, a repeat gets 409 REQUEST_IN_PROGRESS with the Retry-After header, in seconds: wait and send it again to get the answer, at no charge.

When to send again ​

What came backWhat to do
201 with success: false and EXTRACTION_BUSY, EXTRACTION_TIMEOUT or EXTRACTION_UNAVAILABLESend the same request, with the same requestId, later.
409 REQUEST_IN_PROGRESSWait the seconds of Retry-After and send it again.
429Wait the seconds of retryAfter, in the answer's body. A limit per day resets at midnight UTC. See Limits and quotas.
502, 503 or 504, with an HTML pageThe service is restarting: send the same request again after a few seconds.
400, 401, 402, 413 or 422; 201 with DOCUMENT_NOT_RECOGNIZED or IMAGE_NOT_PROCESSABLEDo not send it as it is: fix what the message says. See Errors.

Repeating with the same requestId is safe: a request without data was not charged, and one that already answered is answered again at no charge within 15 minutes.

bash
#!/usr/bin/env bash
# Extract a certificate, handling every answer, with retries that never charge twice.
# Each retry sends the same requestId: a repeat of a finished request gets its
# kept answer at no charge, and a repeat of one still running is told to wait (409).
# Needs DOCSOCR_API_KEY in the environment.
set -uo pipefail

REQUEST_ID="errors-$(date +%s)-$RANDOM" # one id per document, the same on every retry
BODY="{\"imageType\":\"url\",\"imageUrl\":\"https://docsocr.com/samples/certidao-nascimento-exemplo.jpg\",\"requestId\":\"$REQUEST_ID\"}"

for attempt in 0 1 2 3 4; do
  backoff=$((1 << attempt)) # seconds: 1, 2, 4, 8
  # Above the API's 90 s extraction budget; a timeout or a dropped connection is retried
  if ! status=$(curl -sS --max-time 120 -o answer.json -w '%{http_code}' \
    https://api.docsocr.com/api/v1/documents/birth-certificate \
    -H "Authorization: Bearer $DOCSOCR_API_KEY" \
    -H "Content-Type: application/json" \
    --data-binary "$BODY"); then
    sleep "$backoff"
    continue
  fi
  if [ "$status" = 201 ] && grep -q '"success":true' answer.json; then
    cat answer.json
    exit 0
  fi
  error_code=$(grep -o '"errorCode":"[A-Z_]*"' answer.json | cut -d'"' -f4)
  # Wait as long as the answer asks (retryAfter); a limit that resets
  # later, such as a daily quota, is not worth waiting for
  wait=$(grep -o '"retryAfter":[0-9]*' answer.json | cut -d: -f2)
  wait=${wait:-$backoff}
  case "$status:$error_code" in
    # No engine could answer (nothing was charged); still in progress;
    # too many requests; the service restarting
    201:EXTRACTION_BUSY | 201:EXTRACTION_TIMEOUT | 201:EXTRACTION_UNAVAILABLE | 409:* | 429:* | 502:* | 503:* | 504:*)
      if [ "$wait" -le 60 ]; then
        sleep "$wait"
        continue
      fi
      ;;
  esac
  # 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
  # (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
  echo "$status $error_code $(cat answer.json)" >&2
  exit 1
done
echo "No answer after retries: try again later" >&2
exit 1
py
"""Extract a certificate, handling every answer, with retries that never charge twice.

Each retry sends the same requestId: a repeat of a finished request gets its
kept answer at no charge, and a repeat of one still running is told to wait (409).
Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import time
import uuid

import requests

ENDPOINT = "https://api.docsocr.com/api/v1/documents/birth-certificate"
# No engine could answer: nothing was charged, and a retry may succeed
RETRY_LATER = {"EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE"}
# Still in progress, too many requests, the service restarting
RETRY_STATUS = {409, 429, 502, 503, 504}


def extract(image_url, attempts=5):
    request_id = str(uuid.uuid4())  # one id per document, the same on every retry
    for attempt in range(attempts):
        backoff = 2**attempt  # seconds: 1, 2, 4, 8
        try:
            response = requests.post(
                ENDPOINT,
                headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
                json={"imageType": "url", "imageUrl": image_url, "requestId": request_id},
                timeout=120,  # above the API's 90 s extraction budget
            )
        except requests.RequestException:  # a timeout or a dropped connection
            time.sleep(backoff)
            continue
        is_json = response.headers.get("Content-Type", "").startswith("application/json")
        answer = response.json() if is_json else {}
        if response.status_code == 201 and answer.get("success"):
            return answer
        # Wait as long as the answer asks (retryAfter); a limit that resets
        # later, such as a daily quota, is not worth waiting for
        wait = answer.get("retryAfter") or backoff
        retry = answer.get("errorCode") in RETRY_LATER or response.status_code in RETRY_STATUS
        if retry and wait <= 60:
            time.sleep(wait)
            continue
        # 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
        # (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
        raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
    raise SystemExit("No answer after retries: try again later")


answer = extract("https://docsocr.com/samples/certidao-nascimento-exemplo.jpg")
print(answer["data"]["dados_pessoais"]["nome_completo"], "credits:", answer["creditsCharged"])
js
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs Node.js 18+ and DOCSOCR_API_KEY in the environment.
import { randomUUID } from 'node:crypto'
import { setTimeout as sleep } from 'node:timers/promises'

const ENDPOINT = 'https://api.docsocr.com/api/v1/documents/birth-certificate'
// No engine could answer: nothing was charged, and a retry may succeed
const RETRY_LATER = new Set(['EXTRACTION_BUSY', 'EXTRACTION_TIMEOUT', 'EXTRACTION_UNAVAILABLE'])
// Still in progress, too many requests, the service restarting
const RETRY_STATUS = new Set([409, 429, 502, 503, 504])

async function extract(imageUrl, attempts = 5) {
  const requestId = randomUUID() // one id per document, the same on every retry
  for (let attempt = 0; attempt < attempts; attempt++) {
    const backoff = 2 ** attempt // seconds: 1, 2, 4, 8
    let response
    try {
      response = await fetch(ENDPOINT, {
        method: 'POST',
        headers: {
          Authorization: `Bearer ${process.env.DOCSOCR_API_KEY}`,
          'Content-Type': 'application/json',
        },
        body: JSON.stringify({ imageType: 'url', imageUrl, requestId }),
        signal: AbortSignal.timeout(120_000), // above the API's 90 s extraction budget
      })
    } catch {
      await sleep(backoff * 1000) // a timeout or a dropped connection
      continue
    }
    const isJson = (response.headers.get('content-type') ?? '').startsWith('application/json')
    const answer = isJson ? await response.json() : {}
    if (response.status === 201 && answer.success) return answer
    // Wait as long as the answer asks (retryAfter); a limit that resets
    // later, such as a daily quota, is not worth waiting for
    const wait = answer.retryAfter || backoff
    const retry = RETRY_LATER.has(answer.errorCode) || RETRY_STATUS.has(response.status)
    if (retry && wait <= 60) {
      await sleep(wait * 1000)
      continue
    }
    // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
    // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
    throw new Error(`${response.status} ${answer.errorCode ?? ''} ${answer.message ?? answer.error}`)
  }
  throw new Error('No answer after retries: try again later')
}

try {
  const answer = await extract('https://docsocr.com/samples/certidao-nascimento-exemplo.jpg')
  console.log(answer.data.dados_pessoais.nome_completo, 'credits:', answer.creditsCharged)
} catch (error) {
  console.error(error.message)
  process.exit(1)
}
php
<?php
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs the curl extension and DOCSOCR_API_KEY in the environment.

const ENDPOINT = 'https://api.docsocr.com/api/v1/documents/birth-certificate';
// No engine could answer: nothing was charged, and a retry may succeed
const RETRY_LATER = ['EXTRACTION_BUSY', 'EXTRACTION_TIMEOUT', 'EXTRACTION_UNAVAILABLE'];
// Still in progress, too many requests, the service restarting
const RETRY_STATUS = [409, 429, 502, 503, 504];

function extractCertificate(string $imageUrl, int $attempts = 5): array
{
    $requestId = bin2hex(random_bytes(16)); // one id per document, the same on every retry
    for ($attempt = 0; $attempt < $attempts; $attempt++) {
        $backoff = 2 ** $attempt; // seconds: 1, 2, 4, 8
        $request = curl_init(ENDPOINT);
        curl_setopt_array($request, [
            CURLOPT_POST => true,
            CURLOPT_RETURNTRANSFER => true,
            CURLOPT_TIMEOUT => 120, // above the API's 90 s extraction budget
            CURLOPT_HTTPHEADER => [
                'Authorization: Bearer ' . getenv('DOCSOCR_API_KEY'),
                'Content-Type: application/json',
            ],
            CURLOPT_POSTFIELDS => json_encode(['imageType' => 'url', 'imageUrl' => $imageUrl, 'requestId' => $requestId]),
        ]);
        $body = curl_exec($request);
        if ($body === false) { // a timeout or a dropped connection
            sleep($backoff);
            continue;
        }
        $status = curl_getinfo($request, CURLINFO_RESPONSE_CODE);
        $answer = json_decode($body, true) ?? [];
        if ($status === 201 && !empty($answer['success'])) {
            return $answer;
        }
        // Wait as long as the answer asks (retryAfter); a limit that resets
        // later, such as a daily quota, is not worth waiting for
        $wait = (int) ($answer['retryAfter'] ?? $backoff);
        $retry = in_array($answer['errorCode'] ?? '', RETRY_LATER, true) || in_array($status, RETRY_STATUS, true);
        if ($retry && $wait <= 60) {
            sleep($wait);
            continue;
        }
        // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
        // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
        $message = implode('; ', (array) ($answer['message'] ?? $answer['error'] ?? ''));
        throw new RuntimeException("$status " . ($answer['errorCode'] ?? '') . " $message");
    }
    throw new RuntimeException('No answer after retries: try again later');
}

try {
    $answer = extractCertificate('https://docsocr.com/samples/certidao-nascimento-exemplo.jpg');
    echo $answer['data']['dados_pessoais']['nome_completo'], ' credits: ', $answer['creditsCharged'], PHP_EOL;
} catch (RuntimeException $error) {
    fwrite(STDERR, $error->getMessage() . PHP_EOL);
    exit(1);
}
java
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs Java 17+ and DOCSOCR_API_KEY in the environment. Run: java Errors.java
import java.io.IOException;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.Optional;
import java.util.Set;
import java.util.UUID;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class Errors {
    static final String ENDPOINT = "https://api.docsocr.com/api/v1/documents/birth-certificate";
    // No engine could answer: nothing was charged, and a retry may succeed
    static final Set<String> RETRY_LATER = Set.of("EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE");
    // Still in progress, too many requests, the service restarting
    static final Set<Integer> RETRY_STATUS = Set.of(409, 429, 502, 503, 504);
    // Read the answer with your JSON library; this sample only needs two fields
    static final Pattern ERROR_CODE = Pattern.compile("\"errorCode\":\"(\\w+)\"");
    static final Pattern RETRY_AFTER = Pattern.compile("\"retryAfter\":(\\d+)");

    static String extract(String imageUrl, int attempts) throws InterruptedException {
        HttpClient client = HttpClient.newHttpClient();
        String body = """
            {"imageType": "url", "imageUrl": "%s", "requestId": "%s"}"""
            .formatted(imageUrl, UUID.randomUUID()); // one id per document, the same on every retry
        HttpRequest request = HttpRequest.newBuilder(URI.create(ENDPOINT))
            .timeout(Duration.ofSeconds(120)) // above the API's 90 s extraction budget
            .header("Authorization", "Bearer " + System.getenv("DOCSOCR_API_KEY"))
            .header("Content-Type", "application/json")
            .POST(HttpRequest.BodyPublishers.ofString(body))
            .build();

        for (int attempt = 0; attempt < attempts; attempt++) {
            long backoff = 1L << attempt; // seconds: 1, 2, 4, 8
            HttpResponse<String> response;
            try {
                response = client.send(request, HttpResponse.BodyHandlers.ofString());
            } catch (IOException timeoutOrDroppedConnection) {
                Thread.sleep(backoff * 1000);
                continue;
            }
            int status = response.statusCode();
            String answer = response.body();
            if (status == 201 && answer.contains("\"success\":true")) {
                return answer;
            }
            String errorCode = find(ERROR_CODE, answer).orElse("");
            // Wait as long as the answer asks (retryAfter); a limit that resets
            // later, such as a daily quota, is not worth waiting for
            long wait = find(RETRY_AFTER, answer).map(Long::parseLong).orElse(backoff);
            boolean retry = RETRY_LATER.contains(errorCode) || RETRY_STATUS.contains(status);
            if (retry && wait <= 60) {
                Thread.sleep(wait * 1000);
                continue;
            }
            // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
            // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
            throw new IllegalStateException(status + " " + errorCode + " " + answer);
        }
        throw new IllegalStateException("No answer after retries: try again later");
    }

    static Optional<String> find(Pattern pattern, String text) {
        Matcher match = pattern.matcher(text);
        return match.find() ? Optional.of(match.group(1)) : Optional.empty();
    }

    public static void main(String[] args) throws InterruptedException {
        try {
            System.out.println(extract("https://docsocr.com/samples/certidao-nascimento-exemplo.jpg", 5));
        } catch (IllegalStateException error) {
            System.err.println(error.getMessage());
            System.exit(1);
        }
    }
}
csharp
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs .NET 8+ and DOCSOCR_API_KEY in the environment. Run: dotnet run Errors.cs (.NET 10)
using System.Net.Http.Headers;
using System.Text;
using System.Text.Json;
using System.Text.Json.Nodes;

const string Endpoint = "https://api.docsocr.com/api/v1/documents/birth-certificate";
// No engine could answer: nothing was charged, and a retry may succeed
string[] retryLater = ["EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE"];
// Still in progress, too many requests, the service restarting
int[] retryStatus = [409, 429, 502, 503, 504];

using var http = new HttpClient { Timeout = TimeSpan.FromSeconds(120) }; // above the API's 90 s extraction budget
http.DefaultRequestHeaders.Authorization =
    new AuthenticationHeaderValue("Bearer", Environment.GetEnvironmentVariable("DOCSOCR_API_KEY"));
var body = new JsonObject
{
    ["imageType"] = "url",
    ["imageUrl"] = "https://docsocr.com/samples/certidao-nascimento-exemplo.jpg",
    ["requestId"] = Guid.NewGuid().ToString(), // one id per document, the same on every retry
}.ToJsonString();

for (var attempt = 0; attempt < 5; attempt++)
{
    var backoff = TimeSpan.FromSeconds(1 << attempt); // seconds: 1, 2, 4, 8
    using var content = new StringContent(body, Encoding.UTF8, "application/json");
    HttpResponseMessage response;
    try
    {
        response = await http.PostAsync(Endpoint, content);
    }
    catch (Exception error) when (error is HttpRequestException or TaskCanceledException)
    {
        await Task.Delay(backoff); // a timeout or a dropped connection
        continue;
    }
    using (response)
    {
        var status = (int)response.StatusCode;
        var text = await response.Content.ReadAsStringAsync();
        var isJson = response.Content.Headers.ContentType?.MediaType == "application/json";
        using var answer = JsonDocument.Parse(isJson ? text : "{}");
        var root = answer.RootElement;
        var errorCode = root.TryGetProperty("errorCode", out var code) ? code.GetString() : "";

        if (status == 201 && root.TryGetProperty("success", out var ok) && ok.GetBoolean())
        {
            var name = root.GetProperty("data").GetProperty("dados_pessoais").GetProperty("nome_completo").GetString();
            Console.WriteLine($"{name} credits: {root.GetProperty("creditsCharged")}");
            return 0;
        }
        // Wait as long as the answer asks (retryAfter); a limit that resets
        // later, such as a daily quota, is not worth waiting for
        var wait = root.TryGetProperty("retryAfter", out var after) && after.ValueKind == JsonValueKind.Number
            ? TimeSpan.FromSeconds(after.GetInt32())
            : backoff;
        var retry = retryLater.Contains(errorCode) || retryStatus.Contains(status);
        if (retry && wait <= TimeSpan.FromMinutes(1))
        {
            await Task.Delay(wait);
            continue;
        }
        // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
        // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
        Console.Error.WriteLine($"{status} {errorCode} {text}");
        return 1;
    }
}
Console.Error.WriteLine("No answer after retries: try again later");
return 1;

Timeout ​

The API answers an extraction within 90 s. Use a longer timeout in your client, so it does not give up on an answer that is still coming.