IDs de requisição e repetições
O requestId
Cada extração leva um requestId seu: até 128 caracteres, com letras, dígitos, _, -, . e :. Use um por documento, sem dados pessoais nele; um UUID novo para cada certidão serve bem.
Repetir sem pagar duas vezes
O mesmo requestId, com o mesmo arquivo e as mesmas opções, em até 15 minutos da resposta, recebe a mesma resposta sem custo e sem rodar os motores, com o header Idempotent-Replayed: true. Com outro arquivo ou outras opções, o mesmo requestId faz uma nova requisição.
Enquanto a primeira requisição ainda roda, uma repetição recebe 409 REQUEST_IN_PROGRESS com o header Retry-After, em segundos: aguarde e envie de novo para receber a resposta, sem custo.
Quando enviar de novo
| O que voltou | O que fazer |
|---|---|
201 com success: false e EXTRACTION_BUSY, EXTRACTION_TIMEOUT ou EXTRACTION_UNAVAILABLE | Envie a mesma requisição, com o mesmo requestId, mais tarde. |
409 REQUEST_IN_PROGRESS | Aguarde os segundos de Retry-After e envie de novo. |
| 429 | Aguarde os segundos de retryAfter, no corpo da resposta. Um limite por dia volta à meia-noite UTC. Veja Limites e cotas. |
| 502, 503 ou 504, com uma página HTML | O serviço está reiniciando: envie a mesma requisição de novo depois de alguns segundos. |
400, 401, 402, 413 ou 422; 201 com DOCUMENT_NOT_RECOGNIZED ou IMAGE_NOT_PROCESSABLE | Não envie igual: corrija o que a mensagem diz. Veja Erros. |
Repetir com o mesmo requestId é seguro: uma requisição sem dados não foi cobrada, e uma que já respondeu é respondida de novo sem custo em até 15 minutos.
#!/usr/bin/env bash
# Extract a certificate, handling every answer, with retries that never charge twice.
# Each retry sends the same requestId: a repeat of a finished request gets its
# kept answer at no charge, and a repeat of one still running is told to wait (409).
# Needs DOCSOCR_API_KEY in the environment.
set -uo pipefail
REQUEST_ID="errors-$(date +%s)-$RANDOM" # one id per document, the same on every retry
BODY="{\"imageType\":\"url\",\"imageUrl\":\"https://docsocr.com/samples/certidao-nascimento-exemplo.jpg\",\"requestId\":\"$REQUEST_ID\"}"
for attempt in 0 1 2 3 4; do
backoff=$((1 << attempt)) # seconds: 1, 2, 4, 8
# Above the API's 90 s extraction budget; a timeout or a dropped connection is retried
if ! status=$(curl -sS --max-time 120 -o answer.json -w '%{http_code}' \
https://api.docsocr.com/api/v1/documents/birth-certificate \
-H "Authorization: Bearer $DOCSOCR_API_KEY" \
-H "Content-Type: application/json" \
--data-binary "$BODY"); then
sleep "$backoff"
continue
fi
if [ "$status" = 201 ] && grep -q '"success":true' answer.json; then
cat answer.json
exit 0
fi
error_code=$(grep -o '"errorCode":"[A-Z_]*"' answer.json | cut -d'"' -f4)
# Wait as long as the answer asks (retryAfter); a limit that resets
# later, such as a daily quota, is not worth waiting for
wait=$(grep -o '"retryAfter":[0-9]*' answer.json | cut -d: -f2)
wait=${wait:-$backoff}
case "$status:$error_code" in
# No engine could answer (nothing was charged); still in progress;
# too many requests; the service restarting
201:EXTRACTION_BUSY | 201:EXTRACTION_TIMEOUT | 201:EXTRACTION_UNAVAILABLE | 409:* | 429:* | 502:* | 503:* | 504:*)
if [ "$wait" -le 60 ]; then
sleep "$wait"
continue
fi
;;
esac
# 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
# (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
echo "$status $error_code $(cat answer.json)" >&2
exit 1
done
echo "No answer after retries: try again later" >&2
exit 1"""Extract a certificate, handling every answer, with retries that never charge twice.
Each retry sends the same requestId: a repeat of a finished request gets its
kept answer at no charge, and a repeat of one still running is told to wait (409).
Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import time
import uuid
import requests
ENDPOINT = "https://api.docsocr.com/api/v1/documents/birth-certificate"
# No engine could answer: nothing was charged, and a retry may succeed
RETRY_LATER = {"EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE"}
# Still in progress, too many requests, the service restarting
RETRY_STATUS = {409, 429, 502, 503, 504}
def extract(image_url, attempts=5):
request_id = str(uuid.uuid4()) # one id per document, the same on every retry
for attempt in range(attempts):
backoff = 2**attempt # seconds: 1, 2, 4, 8
try:
response = requests.post(
ENDPOINT,
headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
json={"imageType": "url", "imageUrl": image_url, "requestId": request_id},
timeout=120, # above the API's 90 s extraction budget
)
except requests.RequestException: # a timeout or a dropped connection
time.sleep(backoff)
continue
is_json = response.headers.get("Content-Type", "").startswith("application/json")
answer = response.json() if is_json else {}
if response.status_code == 201 and answer.get("success"):
return answer
# Wait as long as the answer asks (retryAfter); a limit that resets
# later, such as a daily quota, is not worth waiting for
wait = answer.get("retryAfter") or backoff
retry = answer.get("errorCode") in RETRY_LATER or response.status_code in RETRY_STATUS
if retry and wait <= 60:
time.sleep(wait)
continue
# 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
# (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
raise SystemExit("No answer after retries: try again later")
answer = extract("https://docsocr.com/samples/certidao-nascimento-exemplo.jpg")
print(answer["data"]["dados_pessoais"]["nome_completo"], "credits:", answer["creditsCharged"])// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs Node.js 18+ and DOCSOCR_API_KEY in the environment.
import { randomUUID } from 'node:crypto'
import { setTimeout as sleep } from 'node:timers/promises'
const ENDPOINT = 'https://api.docsocr.com/api/v1/documents/birth-certificate'
// No engine could answer: nothing was charged, and a retry may succeed
const RETRY_LATER = new Set(['EXTRACTION_BUSY', 'EXTRACTION_TIMEOUT', 'EXTRACTION_UNAVAILABLE'])
// Still in progress, too many requests, the service restarting
const RETRY_STATUS = new Set([409, 429, 502, 503, 504])
async function extract(imageUrl, attempts = 5) {
const requestId = randomUUID() // one id per document, the same on every retry
for (let attempt = 0; attempt < attempts; attempt++) {
const backoff = 2 ** attempt // seconds: 1, 2, 4, 8
let response
try {
response = await fetch(ENDPOINT, {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.DOCSOCR_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ imageType: 'url', imageUrl, requestId }),
signal: AbortSignal.timeout(120_000), // above the API's 90 s extraction budget
})
} catch {
await sleep(backoff * 1000) // a timeout or a dropped connection
continue
}
const isJson = (response.headers.get('content-type') ?? '').startsWith('application/json')
const answer = isJson ? await response.json() : {}
if (response.status === 201 && answer.success) return answer
// Wait as long as the answer asks (retryAfter); a limit that resets
// later, such as a daily quota, is not worth waiting for
const wait = answer.retryAfter || backoff
const retry = RETRY_LATER.has(answer.errorCode) || RETRY_STATUS.has(response.status)
if (retry && wait <= 60) {
await sleep(wait * 1000)
continue
}
// 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
// (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
throw new Error(`${response.status} ${answer.errorCode ?? ''} ${answer.message ?? answer.error}`)
}
throw new Error('No answer after retries: try again later')
}
try {
const answer = await extract('https://docsocr.com/samples/certidao-nascimento-exemplo.jpg')
console.log(answer.data.dados_pessoais.nome_completo, 'credits:', answer.creditsCharged)
} catch (error) {
console.error(error.message)
process.exit(1)
}<?php
// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs the curl extension and DOCSOCR_API_KEY in the environment.
const ENDPOINT = 'https://api.docsocr.com/api/v1/documents/birth-certificate';
// No engine could answer: nothing was charged, and a retry may succeed
const RETRY_LATER = ['EXTRACTION_BUSY', 'EXTRACTION_TIMEOUT', 'EXTRACTION_UNAVAILABLE'];
// Still in progress, too many requests, the service restarting
const RETRY_STATUS = [409, 429, 502, 503, 504];
function extractCertificate(string $imageUrl, int $attempts = 5): array
{
$requestId = bin2hex(random_bytes(16)); // one id per document, the same on every retry
for ($attempt = 0; $attempt < $attempts; $attempt++) {
$backoff = 2 ** $attempt; // seconds: 1, 2, 4, 8
$request = curl_init(ENDPOINT);
curl_setopt_array($request, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_TIMEOUT => 120, // above the API's 90 s extraction budget
CURLOPT_HTTPHEADER => [
'Authorization: Bearer ' . getenv('DOCSOCR_API_KEY'),
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode(['imageType' => 'url', 'imageUrl' => $imageUrl, 'requestId' => $requestId]),
]);
$body = curl_exec($request);
if ($body === false) { // a timeout or a dropped connection
sleep($backoff);
continue;
}
$status = curl_getinfo($request, CURLINFO_RESPONSE_CODE);
$answer = json_decode($body, true) ?? [];
if ($status === 201 && !empty($answer['success'])) {
return $answer;
}
// Wait as long as the answer asks (retryAfter); a limit that resets
// later, such as a daily quota, is not worth waiting for
$wait = (int) ($answer['retryAfter'] ?? $backoff);
$retry = in_array($answer['errorCode'] ?? '', RETRY_LATER, true) || in_array($status, RETRY_STATUS, true);
if ($retry && $wait <= 60) {
sleep($wait);
continue;
}
// 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
// (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
$message = implode('; ', (array) ($answer['message'] ?? $answer['error'] ?? ''));
throw new RuntimeException("$status " . ($answer['errorCode'] ?? '') . " $message");
}
throw new RuntimeException('No answer after retries: try again later');
}
try {
$answer = extractCertificate('https://docsocr.com/samples/certidao-nascimento-exemplo.jpg');
echo $answer['data']['dados_pessoais']['nome_completo'], ' credits: ', $answer['creditsCharged'], PHP_EOL;
} catch (RuntimeException $error) {
fwrite(STDERR, $error->getMessage() . PHP_EOL);
exit(1);
}// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs Java 17+ and DOCSOCR_API_KEY in the environment. Run: java Errors.java
import java.io.IOException;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.Optional;
import java.util.Set;
import java.util.UUID;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
public class Errors {
static final String ENDPOINT = "https://api.docsocr.com/api/v1/documents/birth-certificate";
// No engine could answer: nothing was charged, and a retry may succeed
static final Set<String> RETRY_LATER = Set.of("EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE");
// Still in progress, too many requests, the service restarting
static final Set<Integer> RETRY_STATUS = Set.of(409, 429, 502, 503, 504);
// Read the answer with your JSON library; this sample only needs two fields
static final Pattern ERROR_CODE = Pattern.compile("\"errorCode\":\"(\\w+)\"");
static final Pattern RETRY_AFTER = Pattern.compile("\"retryAfter\":(\\d+)");
static String extract(String imageUrl, int attempts) throws InterruptedException {
HttpClient client = HttpClient.newHttpClient();
String body = """
{"imageType": "url", "imageUrl": "%s", "requestId": "%s"}"""
.formatted(imageUrl, UUID.randomUUID()); // one id per document, the same on every retry
HttpRequest request = HttpRequest.newBuilder(URI.create(ENDPOINT))
.timeout(Duration.ofSeconds(120)) // above the API's 90 s extraction budget
.header("Authorization", "Bearer " + System.getenv("DOCSOCR_API_KEY"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body))
.build();
for (int attempt = 0; attempt < attempts; attempt++) {
long backoff = 1L << attempt; // seconds: 1, 2, 4, 8
HttpResponse<String> response;
try {
response = client.send(request, HttpResponse.BodyHandlers.ofString());
} catch (IOException timeoutOrDroppedConnection) {
Thread.sleep(backoff * 1000);
continue;
}
int status = response.statusCode();
String answer = response.body();
if (status == 201 && answer.contains("\"success\":true")) {
return answer;
}
String errorCode = find(ERROR_CODE, answer).orElse("");
// Wait as long as the answer asks (retryAfter); a limit that resets
// later, such as a daily quota, is not worth waiting for
long wait = find(RETRY_AFTER, answer).map(Long::parseLong).orElse(backoff);
boolean retry = RETRY_LATER.contains(errorCode) || RETRY_STATUS.contains(status);
if (retry && wait <= 60) {
Thread.sleep(wait * 1000);
continue;
}
// 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
// (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
throw new IllegalStateException(status + " " + errorCode + " " + answer);
}
throw new IllegalStateException("No answer after retries: try again later");
}
static Optional<String> find(Pattern pattern, String text) {
Matcher match = pattern.matcher(text);
return match.find() ? Optional.of(match.group(1)) : Optional.empty();
}
public static void main(String[] args) throws InterruptedException {
try {
System.out.println(extract("https://docsocr.com/samples/certidao-nascimento-exemplo.jpg", 5));
} catch (IllegalStateException error) {
System.err.println(error.getMessage());
System.exit(1);
}
}
}// Extract a certificate, handling every answer, with retries that never charge twice.
// Each retry sends the same requestId: a repeat of a finished request gets its
// kept answer at no charge, and a repeat of one still running is told to wait (409).
// Needs .NET 8+ and DOCSOCR_API_KEY in the environment. Run: dotnet run Errors.cs (.NET 10)
using System.Net.Http.Headers;
using System.Text;
using System.Text.Json;
using System.Text.Json.Nodes;
const string Endpoint = "https://api.docsocr.com/api/v1/documents/birth-certificate";
// No engine could answer: nothing was charged, and a retry may succeed
string[] retryLater = ["EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE"];
// Still in progress, too many requests, the service restarting
int[] retryStatus = [409, 429, 502, 503, 504];
using var http = new HttpClient { Timeout = TimeSpan.FromSeconds(120) }; // above the API's 90 s extraction budget
http.DefaultRequestHeaders.Authorization =
new AuthenticationHeaderValue("Bearer", Environment.GetEnvironmentVariable("DOCSOCR_API_KEY"));
var body = new JsonObject
{
["imageType"] = "url",
["imageUrl"] = "https://docsocr.com/samples/certidao-nascimento-exemplo.jpg",
["requestId"] = Guid.NewGuid().ToString(), // one id per document, the same on every retry
}.ToJsonString();
for (var attempt = 0; attempt < 5; attempt++)
{
var backoff = TimeSpan.FromSeconds(1 << attempt); // seconds: 1, 2, 4, 8
using var content = new StringContent(body, Encoding.UTF8, "application/json");
HttpResponseMessage response;
try
{
response = await http.PostAsync(Endpoint, content);
}
catch (Exception error) when (error is HttpRequestException or TaskCanceledException)
{
await Task.Delay(backoff); // a timeout or a dropped connection
continue;
}
using (response)
{
var status = (int)response.StatusCode;
var text = await response.Content.ReadAsStringAsync();
var isJson = response.Content.Headers.ContentType?.MediaType == "application/json";
using var answer = JsonDocument.Parse(isJson ? text : "{}");
var root = answer.RootElement;
var errorCode = root.TryGetProperty("errorCode", out var code) ? code.GetString() : "";
if (status == 201 && root.TryGetProperty("success", out var ok) && ok.GetBoolean())
{
var name = root.GetProperty("data").GetProperty("dados_pessoais").GetProperty("nome_completo").GetString();
Console.WriteLine($"{name} credits: {root.GetProperty("creditsCharged")}");
return 0;
}
// Wait as long as the answer asks (retryAfter); a limit that resets
// later, such as a daily quota, is not worth waiting for
var wait = root.TryGetProperty("retryAfter", out var after) && after.ValueKind == JsonValueKind.Number
? TimeSpan.FromSeconds(after.GetInt32())
: backoff;
var retry = retryLater.Contains(errorCode) || retryStatus.Contains(status);
if (retry && wait <= TimeSpan.FromMinutes(1))
{
await Task.Delay(wait);
continue;
}
// 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
// (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
Console.Error.WriteLine($"{status} {errorCode} {text}");
return 1;
}
}
Console.Error.WriteLine("No answer after retries: try again later");
return 1;Timeout
A API responde a uma extração em até 90 s. Use no seu cliente um timeout maior que esse, para não desistir de uma resposta que ainda vem.