Pular para o conteúdo

IDs de requisição e repetições ​

O requestId ​

Cada extração leva um requestId seu: até 128 caracteres, com letras, dígitos, _, -, . e :. Use um por documento, sem dados pessoais nele; um UUID novo para cada certidão serve bem.

Repetir sem pagar duas vezes ​

O mesmo requestId, com o mesmo arquivo e as mesmas opções, em até 15 minutos da resposta, recebe a mesma resposta sem custo e sem rodar os motores, com o header Idempotent-Replayed: true. Com outro arquivo ou outras opções, o mesmo requestId faz uma nova requisição.

Enquanto a primeira requisição ainda roda, uma repetição recebe 409 REQUEST_IN_PROGRESS com o header Retry-After, em segundos: aguarde e envie de novo para receber a resposta, sem custo.

Quando enviar de novo ​

O que voltouO que fazer
201 com success: false e EXTRACTION_BUSY, EXTRACTION_TIMEOUT ou EXTRACTION_UNAVAILABLEEnvie a mesma requisição, com o mesmo requestId, mais tarde.
409 REQUEST_IN_PROGRESSAguarde os segundos de Retry-After e envie de novo.
429Aguarde os segundos de retryAfter, no corpo da resposta. Um limite por dia volta à meia-noite UTC. Veja Limites e cotas.
502, 503 ou 504, com uma página HTMLO serviço está reiniciando: envie a mesma requisição de novo depois de alguns segundos.
400, 401, 402, 413 ou 422; 201 com DOCUMENT_NOT_RECOGNIZED ou IMAGE_NOT_PROCESSABLENão envie igual: corrija o que a mensagem diz. Veja Erros.

Repetir com o mesmo requestId é seguro: uma requisição sem dados não foi cobrada, e uma que já respondeu é respondida de novo sem custo em até 15 minutos.

bash
#!/usr/bin/env bash
# Extract a certificate, handling every answer, with retries that never charge twice.
# Each retry sends the same requestId: a repeat of a finished request gets its
# kept answer at no charge, and a repeat of one still running is told to wait (409).
# Needs DOCSOCR_API_KEY in the environment.
set -uo pipefail

REQUEST_ID="errors-$(date +%s)-$RANDOM" # one id per document, the same on every retry
BODY="{\"imageType\":\"url\",\"imageUrl\":\"https://docsocr.com/samples/certidao-nascimento-exemplo.jpg\",\"requestId\":\"$REQUEST_ID\"}"

for attempt in 0 1 2 3 4; do
  backoff=$((1 << attempt)) # seconds: 1, 2, 4, 8
  # Above the API's 90 s extraction budget; a timeout or a dropped connection is retried
  if ! status=$(curl -sS --max-time 120 -o answer.json -w '%{http_code}' \
    https://api.docsocr.com/api/v1/documents/birth-certificate \
    -H "Authorization: Bearer $DOCSOCR_API_KEY" \
    -H "Content-Type: application/json" \
    --data-binary "$BODY"); then
    sleep "$backoff"
    continue
  fi
  if [ "$status" = 201 ] && grep -q '"success":true' answer.json; then
    cat answer.json
    exit 0
  fi
  error_code=$(grep -o '"errorCode":"[A-Z_]*"' answer.json | cut -d'"' -f4)
  # Wait as long as the answer asks (retryAfter); a limit that resets
  # later, such as a daily quota, is not worth waiting for
  wait=$(grep -o '"retryAfter":[0-9]*' answer.json | cut -d: -f2)
  wait=${wait:-$backoff}
  case "$status:$error_code" in
    # No engine could answer (nothing was charged); still in progress;
    # too many requests; the service restarting
    201:EXTRACTION_BUSY | 201:EXTRACTION_TIMEOUT | 201:EXTRACTION_UNAVAILABLE | 409:* | 429:* | 502:* | 503:* | 504:*)
      if [ "$wait" -le 60 ]; then
        sleep "$wait"
        continue
      fi
      ;;
  esac
  # 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
  # (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
  echo "$status $error_code $(cat answer.json)" >&2
  exit 1
done
echo "No answer after retries: try again later" >&2
exit 1
py
"""Extract a certificate, handling every answer, with retries that never charge twice.

Each retry sends the same requestId: a repeat of a finished request gets its
kept answer at no charge, and a repeat of one still running is told to wait (409).
Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import time
import uuid

import requests

ENDPOINT = "https://api.docsocr.com/api/v1/documents/birth-certificate"
# No engine could answer: nothing was charged, and a retry may succeed
RETRY_LATER = {"EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE"}
# Still in progress, too many requests, the service restarting
RETRY_STATUS = {409, 429, 502, 503, 504}


def extract(image_url, attempts=5):
    request_id = str(uuid.uuid4())  # one id per document, the same on every retry
    for attempt in range(attempts):
        backoff = 2**attempt  # seconds: 1, 2, 4, 8
        try:
            response = requests.post(
                ENDPOINT,
                headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
                json={"imageType": "url", "imageUrl": image_url, "requestId": request_id},
                timeout=120,  # above the API's 90 s extraction budget
            )
        except requests.RequestException:  # a timeout or a dropped connection
            time.sleep(backoff)
            continue
        is_json = response.headers.get("Content-Type", "").startswith("application/json")
        answer = response.json() if is_json else {}
        if response.status_code == 201 and answer.get("success"):
            return answer
        # Wait as long as the answer asks (retryAfter); a limit that resets
        # later, such as a daily quota, is not worth waiting for
        wait = answer.get("retryAfter") or backoff
        retry = answer.get("errorCode") in RETRY_LATER or response.status_code in RETRY_STATUS
        if retry and wait <= 60:
            time.sleep(wait)
            continue
        # 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
        # (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
        raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
    raise SystemExit("No answer after retries: try again later")


answer = extract("https://docsocr.com/samples/certidao-nascimento-exemplo.jpg")
print(answer["data"]["dados_pessoais"]["nome_completo"], "credits:", answer["creditsCharged"])
js
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs Node.js 18+ and DOCSOCR_API_KEY in the environment.
import { randomUUID } from 'node:crypto'
import { setTimeout as sleep } from 'node:timers/promises'

const ENDPOINT = 'https://api.docsocr.com/api/v1/documents/birth-certificate'
// No engine could answer: nothing was charged, and a retry may succeed
const RETRY_LATER = new Set(['EXTRACTION_BUSY', 'EXTRACTION_TIMEOUT', 'EXTRACTION_UNAVAILABLE'])
// Still in progress, too many requests, the service restarting
const RETRY_STATUS = new Set([409, 429, 502, 503, 504])

async function extract(imageUrl, attempts = 5) {
  const requestId = randomUUID() // one id per document, the same on every retry
  for (let attempt = 0; attempt < attempts; attempt++) {
    const backoff = 2 ** attempt // seconds: 1, 2, 4, 8
    let response
    try {
      response = await fetch(ENDPOINT, {
        method: 'POST',
        headers: {
          Authorization: `Bearer ${process.env.DOCSOCR_API_KEY}`,
          'Content-Type': 'application/json',
        },
        body: JSON.stringify({ imageType: 'url', imageUrl, requestId }),
        signal: AbortSignal.timeout(120_000), // above the API's 90 s extraction budget
      })
    } catch {
      await sleep(backoff * 1000) // a timeout or a dropped connection
      continue
    }
    const isJson = (response.headers.get('content-type') ?? '').startsWith('application/json')
    const answer = isJson ? await response.json() : {}
    if (response.status === 201 && answer.success) return answer
    // Wait as long as the answer asks (retryAfter); a limit that resets
    // later, such as a daily quota, is not worth waiting for
    const wait = answer.retryAfter || backoff
    const retry = RETRY_LATER.has(answer.errorCode) || RETRY_STATUS.has(response.status)
    if (retry && wait <= 60) {
      await sleep(wait * 1000)
      continue
    }
    // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
    // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
    throw new Error(`${response.status} ${answer.errorCode ?? ''} ${answer.message ?? answer.error}`)
  }
  throw new Error('No answer after retries: try again later')
}

try {
  const answer = await extract('https://docsocr.com/samples/certidao-nascimento-exemplo.jpg')
  console.log(answer.data.dados_pessoais.nome_completo, 'credits:', answer.creditsCharged)
} catch (error) {
  console.error(error.message)
  process.exit(1)
}
php
<?php
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs the curl extension and DOCSOCR_API_KEY in the environment.

const ENDPOINT = 'https://api.docsocr.com/api/v1/documents/birth-certificate';
// No engine could answer: nothing was charged, and a retry may succeed
const RETRY_LATER = ['EXTRACTION_BUSY', 'EXTRACTION_TIMEOUT', 'EXTRACTION_UNAVAILABLE'];
// Still in progress, too many requests, the service restarting
const RETRY_STATUS = [409, 429, 502, 503, 504];

function extractCertificate(string $imageUrl, int $attempts = 5): array
{
    $requestId = bin2hex(random_bytes(16)); // one id per document, the same on every retry
    for ($attempt = 0; $attempt < $attempts; $attempt++) {
        $backoff = 2 ** $attempt; // seconds: 1, 2, 4, 8
        $request = curl_init(ENDPOINT);
        curl_setopt_array($request, [
            CURLOPT_POST => true,
            CURLOPT_RETURNTRANSFER => true,
            CURLOPT_TIMEOUT => 120, // above the API's 90 s extraction budget
            CURLOPT_HTTPHEADER => [
                'Authorization: Bearer ' . getenv('DOCSOCR_API_KEY'),
                'Content-Type: application/json',
            ],
            CURLOPT_POSTFIELDS => json_encode(['imageType' => 'url', 'imageUrl' => $imageUrl, 'requestId' => $requestId]),
        ]);
        $body = curl_exec($request);
        if ($body === false) { // a timeout or a dropped connection
            sleep($backoff);
            continue;
        }
        $status = curl_getinfo($request, CURLINFO_RESPONSE_CODE);
        $answer = json_decode($body, true) ?? [];
        if ($status === 201 && !empty($answer['success'])) {
            return $answer;
        }
        // Wait as long as the answer asks (retryAfter); a limit that resets
        // later, such as a daily quota, is not worth waiting for
        $wait = (int) ($answer['retryAfter'] ?? $backoff);
        $retry = in_array($answer['errorCode'] ?? '', RETRY_LATER, true) || in_array($status, RETRY_STATUS, true);
        if ($retry && $wait <= 60) {
            sleep($wait);
            continue;
        }
        // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
        // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
        $message = implode('; ', (array) ($answer['message'] ?? $answer['error'] ?? ''));
        throw new RuntimeException("$status " . ($answer['errorCode'] ?? '') . " $message");
    }
    throw new RuntimeException('No answer after retries: try again later');
}

try {
    $answer = extractCertificate('https://docsocr.com/samples/certidao-nascimento-exemplo.jpg');
    echo $answer['data']['dados_pessoais']['nome_completo'], ' credits: ', $answer['creditsCharged'], PHP_EOL;
} catch (RuntimeException $error) {
    fwrite(STDERR, $error->getMessage() . PHP_EOL);
    exit(1);
}
java
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs Java 17+ and DOCSOCR_API_KEY in the environment. Run: java Errors.java
import java.io.IOException;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.Optional;
import java.util.Set;
import java.util.UUID;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class Errors {
    static final String ENDPOINT = "https://api.docsocr.com/api/v1/documents/birth-certificate";
    // No engine could answer: nothing was charged, and a retry may succeed
    static final Set<String> RETRY_LATER = Set.of("EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE");
    // Still in progress, too many requests, the service restarting
    static final Set<Integer> RETRY_STATUS = Set.of(409, 429, 502, 503, 504);
    // Read the answer with your JSON library; this sample only needs two fields
    static final Pattern ERROR_CODE = Pattern.compile("\"errorCode\":\"(\\w+)\"");
    static final Pattern RETRY_AFTER = Pattern.compile("\"retryAfter\":(\\d+)");

    static String extract(String imageUrl, int attempts) throws InterruptedException {
        HttpClient client = HttpClient.newHttpClient();
        String body = """
            {"imageType": "url", "imageUrl": "%s", "requestId": "%s"}"""
            .formatted(imageUrl, UUID.randomUUID()); // one id per document, the same on every retry
        HttpRequest request = HttpRequest.newBuilder(URI.create(ENDPOINT))
            .timeout(Duration.ofSeconds(120)) // above the API's 90 s extraction budget
            .header("Authorization", "Bearer " + System.getenv("DOCSOCR_API_KEY"))
            .header("Content-Type", "application/json")
            .POST(HttpRequest.BodyPublishers.ofString(body))
            .build();

        for (int attempt = 0; attempt < attempts; attempt++) {
            long backoff = 1L << attempt; // seconds: 1, 2, 4, 8
            HttpResponse<String> response;
            try {
                response = client.send(request, HttpResponse.BodyHandlers.ofString());
            } catch (IOException timeoutOrDroppedConnection) {
                Thread.sleep(backoff * 1000);
                continue;
            }
            int status = response.statusCode();
            String answer = response.body();
            if (status == 201 && answer.contains("\"success\":true")) {
                return answer;
            }
            String errorCode = find(ERROR_CODE, answer).orElse("");
            // Wait as long as the answer asks (retryAfter); a limit that resets
            // later, such as a daily quota, is not worth waiting for
            long wait = find(RETRY_AFTER, answer).map(Long::parseLong).orElse(backoff);
            boolean retry = RETRY_LATER.contains(errorCode) || RETRY_STATUS.contains(status);
            if (retry && wait <= 60) {
                Thread.sleep(wait * 1000);
                continue;
            }
            // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
            // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
            throw new IllegalStateException(status + " " + errorCode + " " + answer);
        }
        throw new IllegalStateException("No answer after retries: try again later");
    }

    static Optional<String> find(Pattern pattern, String text) {
        Matcher match = pattern.matcher(text);
        return match.find() ? Optional.of(match.group(1)) : Optional.empty();
    }

    public static void main(String[] args) throws InterruptedException {
        try {
            System.out.println(extract("https://docsocr.com/samples/certidao-nascimento-exemplo.jpg", 5));
        } catch (IllegalStateException error) {
            System.err.println(error.getMessage());
            System.exit(1);
        }
    }
}
csharp
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs .NET 8+ and DOCSOCR_API_KEY in the environment. Run: dotnet run Errors.cs (.NET 10)
using System.Net.Http.Headers;
using System.Text;
using System.Text.Json;
using System.Text.Json.Nodes;

const string Endpoint = "https://api.docsocr.com/api/v1/documents/birth-certificate";
// No engine could answer: nothing was charged, and a retry may succeed
string[] retryLater = ["EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE"];
// Still in progress, too many requests, the service restarting
int[] retryStatus = [409, 429, 502, 503, 504];

using var http = new HttpClient { Timeout = TimeSpan.FromSeconds(120) }; // above the API's 90 s extraction budget
http.DefaultRequestHeaders.Authorization =
    new AuthenticationHeaderValue("Bearer", Environment.GetEnvironmentVariable("DOCSOCR_API_KEY"));
var body = new JsonObject
{
    ["imageType"] = "url",
    ["imageUrl"] = "https://docsocr.com/samples/certidao-nascimento-exemplo.jpg",
    ["requestId"] = Guid.NewGuid().ToString(), // one id per document, the same on every retry
}.ToJsonString();

for (var attempt = 0; attempt < 5; attempt++)
{
    var backoff = TimeSpan.FromSeconds(1 << attempt); // seconds: 1, 2, 4, 8
    using var content = new StringContent(body, Encoding.UTF8, "application/json");
    HttpResponseMessage response;
    try
    {
        response = await http.PostAsync(Endpoint, content);
    }
    catch (Exception error) when (error is HttpRequestException or TaskCanceledException)
    {
        await Task.Delay(backoff); // a timeout or a dropped connection
        continue;
    }
    using (response)
    {
        var status = (int)response.StatusCode;
        var text = await response.Content.ReadAsStringAsync();
        var isJson = response.Content.Headers.ContentType?.MediaType == "application/json";
        using var answer = JsonDocument.Parse(isJson ? text : "{}");
        var root = answer.RootElement;
        var errorCode = root.TryGetProperty("errorCode", out var code) ? code.GetString() : "";

        if (status == 201 && root.TryGetProperty("success", out var ok) && ok.GetBoolean())
        {
            var name = root.GetProperty("data").GetProperty("dados_pessoais").GetProperty("nome_completo").GetString();
            Console.WriteLine($"{name} credits: {root.GetProperty("creditsCharged")}");
            return 0;
        }
        // Wait as long as the answer asks (retryAfter); a limit that resets
        // later, such as a daily quota, is not worth waiting for
        var wait = root.TryGetProperty("retryAfter", out var after) && after.ValueKind == JsonValueKind.Number
            ? TimeSpan.FromSeconds(after.GetInt32())
            : backoff;
        var retry = retryLater.Contains(errorCode) || retryStatus.Contains(status);
        if (retry && wait <= TimeSpan.FromMinutes(1))
        {
            await Task.Delay(wait);
            continue;
        }
        // 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
        // (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
        Console.Error.WriteLine($"{status} {errorCode} {text}");
        return 1;
    }
}
Console.Error.WriteLine("No answer after retries: try again later");
return 1;

Timeout ​

A API responde a uma extração em até 90 s. Use no seu cliente um timeout maior que esse, para não desistir de uma resposta que ainda vem.