Reading Shared Indexes
Gateways that take part in Index Sharing publish signed indexes that locate data items inside their bundles. Gateways use them automatically, but nothing about them is gateway-only: any client can fetch a publication, verify it, and look up an ID itself. This page shows how. To have your own gateway subscribe or publish, see Index Sharing.
The Publication
A publishing gateway serves one JSON publication at /ar-io/indexes:
{
"version": 1,
"publisher": "34LYvMptiDvBP5sqfh1oAd6Q4qFsy4PWaZ1HTFmML7h5",
"sequence": 12,
"issuedAt": "2026-10-01T12:00:00.000Z",
"expiresAt": "2026-10-02T12:00:00.000Z",
"indexes": [
{
"name": "root-tx-index",
"kind": "cdb64-root-tx",
"bands": [
{
"id": "b1-h1950000-tip-20260917",
"heightRange": [1950000, null],
"files": [
{ "name": "manifest.json", "size": 41233, "sha256": "…" },
{ "name": "00.cdb", "size": 7012345, "sha256": "…" }
]
}
]
}
],
"signature": { "alg": "ed25519", "keyId": "34LYv…", "sig": "…" }
}publisher is the gateway's wallet, signature.keyId its observer address, and signature.sig a base64 Ed25519 signature. The publication may carry fields not shown here, and later versions may add more. Keep them: the signature covers them.
Verify It
Get the Publisher's Registered Observer Address
Look up the publisher's wallet in the gateway registry: with the ar.io SDK, or from any gateway's /ar-io/peers. The publication must be signed by that gateway's observerAddress. A signature from any other key proves only that somebody signed something.
Check the Signature
The signed message is the prefix ar-io-index-publication/v1 and a newline, followed by the RFC 8785 canonical JSON of the publication with signature removed. This works in Node.js 20+ and current browsers, with one dependency (npm install json-canonicalize):
import { canonicalize } from "json-canonicalize";
const B58 = "123456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz";
function base58Decode(s) {
let n = 0n;
for (const c of s) n = n * 58n + BigInt(B58.indexOf(c));
const bytes = [];
while (n > 0n) {
bytes.unshift(Number(n % 256n));
n /= 256n;
}
for (const c of s) {
if (c !== "1") break;
bytes.unshift(0);
}
return new Uint8Array(bytes);
}
async function verifyPublication(doc, observerAddress) {
const { signature, ...unsigned } = doc;
if (signature?.alg !== "ed25519") throw new Error("unknown algorithm");
if (signature.keyId !== observerAddress) throw new Error("not the registered key");
const key = await crypto.subtle.importKey(
"raw",
base58Decode(signature.keyId),
{ name: "Ed25519" },
false,
["verify"],
);
const message = new TextEncoder().encode(
"ar-io-index-publication/v1\n" + canonicalize(unsigned),
);
const sig = Uint8Array.from(atob(signature.sig), (c) => c.charCodeAt(0));
return crypto.subtle.verify("Ed25519", key, sig, message);
}
const doc = await fetch("https://turbo-gateway.com/ar-io/indexes").then((r) => r.json());
console.log(await verifyPublication(doc, "<observerAddress from the registry>"));Use a real RFC 8785 library. A key-sorted JSON.stringify gives different
bytes for some numbers and keys, and the signature will not verify.
Check It Is Current
Remember the highest sequence you have accepted from each publisher, and refuse a lower one: a cache or mirror could serve an older publication. If expiresAt has passed, the publisher has stopped signing; its bands are still valid, but nothing newer is coming.
Fetch a Band File
Fetch files by their SHA-256. The address can't change meaning, so it is safe to cache and can come from any server that has the file:
async function fetchFile(gateway, file) {
const res = await fetch(`${gateway}/ar-io/indexes/blob/${file.sha256}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
const digest = await crypto.subtle.digest("SHA-256", bytes);
const hex = [...new Uint8Array(digest)].map((b) => b.toString(16).padStart(2, "0")).join("");
if (bytes.length !== file.size || hex !== file.sha256) throw new Error("file does not match");
return bytes;
}Range requests are supported, which is how a large file resumes. Peers seeding a band over BitTorrent are not metered; see Bands as torrents. A 503 with Retry-After means the publisher is part way through replacing a band: fetch the publication again after the delay.
Bands as Torrents
A publisher that runs a torrent engine adds an optional torrent entry to each band it seeds:
"torrent": {
"infohashV1": "<40 hex characters>",
"infohashV2": "<64 hex characters>",
"magnet": "magnet:?xt=urn:btih:…",
"torrentUrl": "/ar-io/indexes/torrents/<infohashV1>.torrent"
}| Field | Meaning |
|---|---|
infohashV1 | The v1 infohash, 40 hex characters |
infohashV2 | The v2 infohash, 64 hex characters. Torrents are hybrid v1 + v2 |
magnet | A magnet link for the torrent |
torrentUrl | Where the .torrent file is served. On an ar.io gateway it is addressed by the v1 infohash, so a band rebuilt under the same id gets a new URL |
The entry is absent when the publisher runs no engine; the HTTP routes always work. The torrent route can answer 404 until the publisher has built a band's torrent, which is normal.
The torrent name is derived from content, not the band id: the first 16 hex characters of the SHA-256 over one line per file, <name>\0<size>\0<sha256 hex>\n, with files in bytewise name order. So publishers of the same bytes share one infohash and one swarm. It is also the <torrent name> in the WebSeed route, /ar-io/indexes/webseed/<torrent name>/<file>, which serves band files to torrent clients and is metered like the blob route.
Only the info dictionary is signed. The infohashes cover the torrent's info dictionary (file names, sizes and piece hashes) and nothing else. Trackers and WebSeeds in a .torrent file are outside it and unsigned. The gateway's own subscriber checks a .torrent against the signed infohashes. It checks that the file list is exactly the band's signed files and sizes (plus BEP 47 pad files), drops every WebSeed it names, and hashes every downloaded file against its signed SHA-256, as over HTTP. A client of its own should do the same.
Pull a Whole Index with Any BitTorrent Client
To mirror a publisher's full index, for a pipeline or an agent, use the magnet link or the .torrent file with any BitTorrent client. Peers aren't metered, so this is the cheapest way to take everything. Before using a .torrent, check its infohash matches the signed infohashV1, since the file itself is unsigned. After the download, check every file's size and SHA-256 against the signed publication, exactly as for an HTTP fetch: the infohash protects the pieces, but only the publication says these are the right files. Seed it afterwards if you can; that is what keeps the swarm fast. Downloading Indexes with BitTorrent walks through it with qBittorrent and aria2.
Look Up One ID
A cdb64-root-tx band is a partitioned CDB64 index. To find the root transaction of one data item:
- Fetch and check the band's
manifest.json. - Pick the partition for the ID's first byte: the manifest lists partitions by two-character hex prefix.
- Fetch and check that one partition, up to about 30 MB.
- Look the 32-byte ID up in it. The value is MessagePack, holding the root transaction ID and, when known, the item's byte offsets.
Bands may overlap. A publisher should never have two bands disagree about an item, since an item has one location, so search them newest first and stop at the first match.
The reference client does all of this in about 130 lines of Python, with the CDB64 reader and MessagePack decoder written out.
Query a Parquet Dataset
Release Requirement: parquet-l1 bands come with gateway Release 85,
which is not released yet. Some gateways, such as turbo-gateway.com, already
publish them from pre-release builds.
A cdb64-root-tx band answers one question quickly: where is this ID. A parquet-l1 band answers a different kind: how many, how large, which, over the whole chain. It holds Arweave's base layer as Apache Parquet, partitioned into bands by block height, so any engine that reads Parquet can query it. No band files to download, no import into a database, and no API to learn. The examples use DuckDB, which installs its own httpfs and json extensions on first use.
See What a Gateway Publishes
Not every gateway publishes every dataset. parquet-l1 is opt-in, so check first:
curl -s https://turbo-gateway.com/ar-io/indexes \
| jq '.indexes[] | {name, kind, bands: (.bands | length)}'{ "name": "root-tx-index", "kind": "cdb64-root-tx", "bands": 5 }
{ "name": "parquet-l1", "kind": "parquet-l1", "bands": 24 }name is the dataset you want; kind tells you how to read it. They are not the same thing, and for root-tx-index they differ.
The Tables
Each band directory holds one Parquet file per table, plus a band.json with the band's heights, row counts and a schema version:
| File | One row per | Key columns |
|---|---|---|
blocks.parquet | block | height, indep_hash, previous_block, block_size, weave_size, tx_root, hash_list_merkle |
transactions.parquet | transaction | id, height, data_size, data_root, owner_address, target, quantity, reward, content_type, format, offset |
tags.parquet | tag | id, height, tag_index, tag_name, tag_value |
block_transactions.parquet | transaction in a block | height, id, block_transaction_index |
wallets.parquet | wallet | address, public_modulus |
IDs, addresses, hashes and tag names and values are stored as raw bytes (BLOB), not text. Base64url is a presentation format, so convert when you want to read one.
Lookup Files
Release Requirement: lookup files (layout l1-3) come with Release 85,
which is not released yet.
From layout l1-3 (the schema in a band's publication metadata, and in its band.json), a band also holds three lookup files. Parquet has no index, so finding one transaction otherwise means scanning every band's id column. A lookup file is derived from the band's tables and sorted by a key, in row groups of 16,384 rows. Parquet keeps each row group's min and max, so a reader holding a key reads the footer, then the one or two row groups that can hold it.
| File | One row per | Columns |
|---|---|---|
lookup_tx_id.parquet | transaction | id8, height |
lookup_wallet.parquet | transaction an address signed (role 0) or received (role 1) | addr8, role, height, data_size |
lookup_tag.parquet | distinct tag name and value | name8, val8, name, value, txs, first_height, last_height |
The keys are unsigned 64-bit integers made one of two ways, which any client can reproduce:
prefix64: the first 8 bytes, big-endian, zero-padded on the right if shorter. For IDs and addresses. In DuckDB:('0x' || rpad(left(hex(x), 16), 16, '0'))::UBIGINT.sha256_64:prefix64of the value's SHA-256. For tag names and values. In DuckDB:('0x' || left(sha256(x), 16))::UBIGINT.
| Encoding | Input | Output |
|---|---|---|
prefix64 | the 32 bytes of ID O048e9pT5nX1CPrMjGC1y1dWdtd3AChFX27hoRsVIdA | 0x3b4e3c7bda53e675 |
prefix64 | the single byte 0xab | 0xab00000000000000 |
sha256_64 | App-Name | 0xbf6cc2a967f23a82 |
sha256_64 | ArDrive-App | 0xa2c30101e8045f65 |
A key is a pointer: two values can share one, so always finish in the table the key points into, as the examples below do. data_size and txs are carried so a wallet's bytes stored and a tag's count need no table read. A lookup file is listed and signed in the publication like any band file. In a gateway checkout, ./tools/ar-io-node index-l1-verify --bands-dir (Release 85) checks that it holds exactly what its band's tables give.
parquet-l1 is the base layer only: Arweave transactions, not the data
items bundled inside them. For items inside bundles, use
root-tx-index or the gateway's GraphQL.
Query It in Place, Downloading Nothing
The byte routes serve HTTP range requests, and Parquet is designed to be read that way: the reader fetches the footer, decides which row groups it needs, and fetches only those. So a query against a dataset on a gateway moves a small fraction of it.
Only a publisher serves the byte routes. A gateway that subscribes to a dataset answers 404 for them, so point queries at the publisher named in the publication.
This is a complete script. DuckDB is the only dependency, and it reads the publication itself to find the bands:
INSTALL httpfs; LOAD httpfs; INSTALL json; LOAD json;
SET VARIABLE gateway = 'https://turbo-gateway.com';
SET VARIABLE txs = (
SELECT list(getvariable('gateway') || '/ar-io/indexes/parquet-l1/' || b.id || '/transactions.parquet')
FROM (SELECT unnest(ix.bands) AS b
FROM (SELECT unnest(indexes) AS ix
FROM read_json(getvariable('gateway') || '/ar-io/indexes'))
WHERE ix.name = 'parquet-l1')
);
SELECT count(*) AS transactions
FROM read_parquet(getvariable('txs'))
WHERE height BETWEEN 1500000 AND 1500099;Swap the last statement for whatever you want to ask. The largest transactions ever posted to the base layer, with their IDs converted back to base64url:
SELECT
rtrim(replace(replace(to_base64(id), '+', '-'), '/', '_'), '=') AS tx_id,
height,
round(data_size / power(2, 30), 1) AS gib
FROM read_parquet(getvariable('txs'))
ORDER BY data_size DESC
LIMIT 3;┌─────────────────────────────────────────────┬─────────┬───────┐
│ tx_id │ height │ gib │
├─────────────────────────────────────────────┼─────────┼───────┤
│ oWRzBr3KHhULAL-s5ULeXac1mb_WQOX5uFBRea16iRI │ 1992471 │ 621.9 │
│ 7vg2832WFsisEcBr1oBQ8ldc4EGOkjQdwW46hDvJsOs │ 1989512 │ 159.6 │
│ SFUs16XzaWFhWvWOR55hb-5X081r2LeWWAHq4HWUt00 │ 1372669 │ 33.7 │
└─────────────────────────────────────────────┴─────────┴───────┘Tags and content types work the same way. Both columns are bytes, so cast them:
SELECT CAST(tag_name AS VARCHAR) AS tag, count(*) AS n
FROM read_parquet(getvariable('gateway') || '/ar-io/indexes/parquet-l1/<band-id>/tags.parquet')
GROUP BY 1 ORDER BY n DESC LIMIT 5;Find One Transaction
Release Requirement: this example needs l1-3 bands, which come with
Release 85, not released yet.
Point lookups are what a scan is bad at: finding one ID means reading every band's id column. The lookup files turn it into two small reads, first the height from lookup_tx_id, then the rows at that height. This is a complete script:
INSTALL httpfs; LOAD httpfs; INSTALL json; LOAD json;
SET VARIABLE gateway = 'https://vilenarios.com';
SET VARIABLE tx = 'dwXPQC8J1u2M8nr5uP-rDgi5JVvpQ5MrjMNRjleUyzk';
-- Every band's file prefix and layout, from the publication.
CREATE TEMP TABLE bands AS
SELECT getvariable('gateway') || '/ar-io/indexes/parquet-l1/' || b.id || '/' AS dir,
b.metadata.schema AS schema
FROM (SELECT unnest(ix.bands) AS b
FROM (SELECT unnest(indexes) AS ix
FROM read_json(getvariable('gateway') || '/ar-io/indexes'))
WHERE ix.name = 'parquet-l1');
SET VARIABLE lookups = (SELECT list(dir || 'lookup_tx_id.parquet') FROM bands WHERE schema = 'l1-3');
SET VARIABLE txs = (SELECT list(dir || 'transactions.parquet') FROM bands);
-- The ID as bytes, and its key: the first 8 bytes as an unsigned integer.
SET VARIABLE id = from_base64(rpad(replace(replace(getvariable('tx'), '-', '+'), '_', '/'), 44, '='));
SET VARIABLE id8 = ('0x' || left(hex(getvariable('id')), 16))::UBIGINT;
-- 1. The lookup files give the height (a key can match more than one).
SET VARIABLE heights = (
SELECT list(DISTINCT height) FROM read_parquet(getvariable('lookups'))
WHERE id8 = getvariable('id8')
);
-- 2. The transaction, from the rows at those heights.
SELECT height, data_size, CAST(content_type AS VARCHAR) AS content_type
FROM read_parquet(getvariable('txs'))
WHERE height BETWEEN list_min(getvariable('heights')) AND list_max(getvariable('heights'))
AND id = getvariable('id');┌─────────┬───────────┬──────────────┐
│ height │ data_size │ content_type │
├─────────┼───────────┼──────────────┤
│ 1150000 │ 36739084 │ NULL │
└─────────┴───────────┴──────────────┘Measured against a gateway publishing the whole chain at l1-3 (24 bands, heights 0 to 2,017,885):
| Finding one transaction | Bytes fetched | Time |
|---|---|---|
| With the lookup files (the script above) | 12.9 MB | 0.7 s |
Scanning id across every band | 2.6 GB | 26 s |
The last step matches the full ID, not the key, so two IDs that share a key can't be confused. The height filter is what keeps it cheap: each band's footer rules out every row group but the one or two holding that height. A band published before layout l1-3 has no lookup files, so a transaction in it is found only by a scan, as in the examples above.
What Actually Crosses the Network
The cost depends on the query, not on the size of the dataset. A count(*) over the whole dataset reads about 4 MB: one footer per band, then one row group. An aggregate over a whole column, such as sum(data_size), reads that column. Filtering on height is the cheapest filter, since bands are partitioned by it and row groups are ordered within a band.
This depends on range requests working end to end. A reverse proxy with a
cache zone in front of a gateway can answer a range request with the whole
file instead, which turns a 4 MB query into hundreds of megabytes. It happens
when the gateway meters its byte routes, because the responses are then
private, nothing can be cached, and the proxy has no stored object to
answer a range from. If a gateway behaves that way, its operator needs
the uncached location block.
Check with curl -s -o /dev/null -w '%{http_code} %{size_download}' -r 0-99 <url>,
which should print 206 100.
Verify Before You Trust It
Querying over HTTP means trusting the server for the bytes it streams you. If the answer matters, verify the publication first, then fetch the band files you need by digest and query them locally. Two things make this cheaper than it sounds:
- Every band file has a signed SHA-256 in the publication, so a local copy is checkable.
Repr-Digestis co-signable, so even a206carries a signature over the digest of the whole file it came from. A client can confirm which file it is reading bytes out of without fetching all of it.
A parquet-l1 dataset is also checkable against Arweave itself, independently of the publisher: block linkage, hash_list_merkle, tx_root and weave accounting can all be recomputed from the rows. In a gateway checkout, ./tools/ar-io-node index-l1-verify --bands-dir (Release 85) does that over a local copy.
Metering
The publication and .torrent files are free. The byte routes are metered like data, so an analytical client is a paying or allowlisted client, with a free allowance first. An accidental full-file download can spend that allowance in one request and the next range request returns 402; that is metering, not a fault. For sustained or heavy work, take the whole dataset over BitTorrent, which is not metered, and query it locally:
SELECT count(*) FROM read_parquet('parquet-l1/*/transactions.parquet')
WHERE height BETWEEN 1000000 AND 1000100;This or GraphQL?
Both read a gateway's index, and they are good at different things.
| Use | When |
|---|---|
| GraphQL | Anything that includes data items inside bundles; anything near the chain tip; filtering on several fields at once; a browser or a small client with no query engine |
parquet-l1 | Counting, summing, grouping or joining over the whole base layer; bulk export with no pagination; reproducible analysis on a pinned, verifiable snapshot; anything where you do not want a gateway in the query path. On l1-3 bands, also one transaction by ID, a wallet's transactions, or a tag value's count, through the lookup files |
GraphQL cannot aggregate: there is no count or sum, and a large result has to be paged. On an l1-3 band, a point lookup is two small reads through the lookup files; on a band of an older layout, it is a scan of every band's id column. A published band set also lags the tip by its publishing cadence.
How is this guide?