Community Archive logo
Agent-friendly documentation

Build with the archive

Download the daily data export, query public records through the API, or give an agent one canonical starting point.

Point an agent here

https://www.community-archive.org/llms.txt

This plain-text index includes the data dump, API credentials, query examples, schema notes, and deeper documentation links.

Choose an access path

Choose the policy-safe surface that fits your query.

Bulk Parquet export

Download tweets and profiles for bulk analysis. Each daily export checks current membership and opt-outs before publication.

One raw archive

Signed-in owners can download their own processed archive JSON while current PostgreSQL policy permits it.

https://www.community-archive.org/api/archive/<username>

Bulk Parquet export

The daily package contains enriched tweets.parquet, separate profiles.parquet, and a manifest.json with row counts, schemas, and checksums. It includes eligible members and applies current consent to referenced authors too.

Use the latest release link below to find the most recent successful export. GitHub lists the download links; the files stay in consent-managed storage. Superseded packages are removed, so bookmark the latest release rather than an individual Parquet URL. Downloads may be temporarily withdrawn when consent changes.

What can I build or discover?

  • Find a remembered quote, reply, link, or old discussion.
  • Revisit your own themes, collaborators, and changing interests.
  • Follow emerging ideas with Trends, Bangers, and the Digest.
  • Trace people, projects, events, and intellectual lineages.
  • Build a personal canon, visualization, or research dataset.
  • Analyze public interaction networks and communities.
More examples and starting points

API quickstart

The API is PostgREST served by Supabase. Send the public anon key in both the apikey and Authorizationheaders. The anon key is intentionally public; never use a service-role key in client code.

cURL

export CA_API_URL='https://fabxmporizzqflnftavs.supabase.co'
export CA_ANON_KEY='eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6ImZhYnhtcG9yaXp6cWZsbmZ0YXZzIiwicm9sZSI6ImFub24iLCJpYXQiOjE3MjIyNDQ5MTIsImV4cCI6MjAzNzgyMDkxMn0.UIEJiUNkLsW28tBHmG-RQDW-I5JNlJLt62CSk9D_qG8'

curl --get "$CA_API_URL/rest/v1/enriched_tweets" \
  -H "apikey: $CA_ANON_KEY" \
  -H "Authorization: Bearer $CA_ANON_KEY" \
  --data-urlencode "select=tweet_id,username,created_at,full_text" \
  --data-urlencode "username=ilike.defenderofbasic" \
  --data-urlencode "order=created_at.desc" \
  --data-urlencode "limit=5"

JavaScript

import { createClient } from '@supabase/supabase-js'

const supabase = createClient(
  'https://fabxmporizzqflnftavs.supabase.co',
  process.env.CA_ANON_KEY,
)

const { data, error } = await supabase
  .from('enriched_tweets')
  .select('tweet_id, username, created_at, full_text')
  .ilike('username', 'defenderofbasic')
  .order('created_at', { ascending: false })
  .limit(5)

if (error) throw error
console.log(data)

Useful resources

  • enriched_tweets — joined tweet and account data
  • tweets — core tweet records
  • all_account — accounts and usernames
  • all_profile — bios and profile media
  • user_directory — directory-ready account data

Query rules

  • Username casing is preserved; use ilike when the original casing is unknown.
  • Treat Twitter IDs as strings, not numbers.
  • Request only the columns you need with select.
  • Paginate large results; responses are capped at 1,000 rows.
  • Use a stable order when paging through changing data.

Deeper docs

The repository contains schema guidance and complete examples.