| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
All checks were successful
Deploy / build-deploy (push) Successful in 7m30s
The AV1-only publish lasted four hours in the wild: the first viewer on
mobile Safari got "bad media error" from the raw file, because a post's
link is fetched raw — Lemmy apps and browsers play that exact URL, with
no <source> negotiation in front of it. The previous commit's note
("--raw ... if that audience matters") had it backwards: the audience
that cannot play AV1 is not a per-post judgement call, it is whoever
happens to open the thread on an iPhone.
So publish-media.sh now emits two encodings and points the post at the
compatible one:
* Every video transcode also produces <hash>.h264.mp4 (x264 crf 23,
same denoise, same frames), uploaded beside the AV1 under the AV1's
hash — the poster's sibling-naming trick, reused, so fetch-media.sh
finds it by name with nothing to look up.
* The printed URL to paste into the post is the H.264 one. pict-rs
can thumbnail it too, so instance thumbnails come back as a bonus.
fetch-media.sh adopting an own-origin .h264.mp4 URL swaps the AV1 back
in as the page's primary when it is on the mount, and records the H.264
as `fallback` in posts.json. The renderer turns a non-empty fallback
into a <source> pair: the AV1 first with an explicit codecs parameter —
both files are video/mp4, so the parameter is the only thing that lets
a non-AV1 browser skip to the file it can play — and the H.264 second.
Browsers with AV1 keep downloading the small file; Safari before 17
gets one that plays instead of an element that will not.
e2e gains the matching conditional gate: a page offering an av01
<source> must offer an .h264.mp4 one, so an AV1 video can never again
ship without its fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
| .. | ||
| Caddyfile.example | ||
| catcrafts-analytics | ||
| catcrafts-analytics.service | ||
| catcrafts-analytics.timer | ||
| catcrafts-deploy.path | ||
| catcrafts-deploy.service | ||
| catcrafts-server.service | ||
| goaccess-browsers.list | ||
| README.md | ||
Deploying catcrafts.net
Two artifacts come out of CI and land in two different places, on purpose.
| goes to | served by | |
|---|---|---|
| wasm bundle + static assets | /srv/catcrafts.net |
Caddy file_server |
catcrafts-server + content/ |
/srv/catcrafts-app |
itself, on 127.0.0.1:8081 |
They are kept apart because the web root is publicly served and mirrored with
rsync --delete on every push. Anything runtime-owned that lives there is
both published to the internet and destroyed on the next deploy — which matters
now for the content files and matters a great deal once the shop has a database
and bank credentials.
Runtime state goes in a third place, /var/lib/catcrafts, created by the
service's StateDirectory=. Today that is orders.jsonl (the order event log)
and bunq-state.json (the bunq session context, including the client RSA key
the server generates on first contact).
Two things on this box cannot be regenerated. Everything else — the wasm bundle, the content, the binary — comes back from a rebuild.
The first is orders.jsonl. It is the ledger: every order and every status
transition, append-only, and the audit trail the tax records lean on. Back it
up:
install -d -m 0700 /var/backups/catcrafts
cp /var/lib/catcrafts/orders.jsonl \
/var/backups/catcrafts/orders-$(date -u +%F).jsonl
It contains names, addresses and email addresses, so it is personal data: keep
it 0600, keep it off the web root, and encrypt it before it leaves the machine.
(bunq-state.json is deliberately NOT worth backing up: delete it and the
server re-onboards from the API key on the next start.)
The second is /srv/catcrafts-app/media — the mirrored post images and screen
recordings. Usually reproducible from content/posts.json, but not if a source
instance has deleted the file since, and for these posts the media is the
content. It lives on the app mount rather than the web root specifically so
rsync --delete cannot reach it, and CI only ever adds to it. Worth including in
the same backup.
Why the backend speaks plaintext HTTP/1.1
Caddy terminates TLS and reverse-proxies to loopback. The backend uses
Crafter.Network's ListenerHTTP1 rather than its HTTP/3 ListenerHTTP for one
concrete reason: Caddy cannot reverse_proxy to an h3 upstream. Port 8081
must never be exposed directly.
One-time host setup
# 1. service account, no login, no home
useradd --system --no-create-home --shell /usr/sbin/nologin catcrafts
# 2. directories
mkdir -p /srv/catcrafts.net /srv/catcrafts-app
chown catcrafts:catcrafts /srv/catcrafts-app
# The web root stays writable by whatever uid the runner container uses.
# 3. units
cp deploy/catcrafts-server.service /etc/systemd/system/
cp deploy/catcrafts-deploy.service /etc/systemd/system/
cp deploy/catcrafts-deploy.path /etc/systemd/system/
systemctl daemon-reload
systemctl enable --now catcrafts-deploy.path
# catcrafts-server is started by the first deploy; enable it so it survives a
# reboot:
systemctl enable catcrafts-server
# 4. Caddy — merge deploy/Caddyfile.example into your site config
caddy validate --config /etc/caddy/Caddyfile
systemctl reload caddy
Runner configuration
Both mounts are required. In the Forgejo runner's config.yaml:
container:
options: "-v /srv/catcrafts.net:/deploy -v /srv/catcrafts-app:/deploy-app"
The deploy steps fail with an explanatory message if either is missing, rather than silently succeeding into the container's own filesystem.
How the restart happens
CI runs inside a container and has no access to the host's systemd. Rather
than give the runner host root or a polkit rule — both of which grant far more
authority than "restart one service" needs — the last deploy step writes
/srv/catcrafts-app/.deploy-stamp, and catcrafts-deploy.path on the host
reacts.
catcrafts-deploy.service runs catcrafts-server --selftest as ExecStartPre
before restarting. That self-test exercises the HTML-escaping and JSON layers
that generate every byte of markup the site emits, so a broken binary leaves
the working server running instead of replacing it.
The binary is installed as catcrafts-server.new and mvd into place, so a
request arriving mid-copy never hits a truncated executable.
If the backend is down
Caddy's handle_errors fallback serves the static index.html,
so the site degrades to the client-rendered wasm app rather than showing a 502.
Content still renders; what is lost is server-side rendering and real status
codes — an unknown path becomes a soft 404 again until the backend returns.
Posts and their media
content/posts-sources.json holds the account and a community allowlist.
Only listed communities are fetched, so joining a new one does not silently
publish it to the site — add a line first.
tools/publish-media.sh FILE # BEFORE posting: uploads to /media, prints the URL
tools/fetch-posts.sh # writes content/posts.json
tools/fetch-media.sh # mirrors the media, rewrites posts.json to /media/... paths
Order matters: the second reads what the first wrote. CI runs both before the
build. Neither fails the build on a network error — a fediverse outage leaves the
previous posts.json in place, and a single failed download leaves that one entry
pointing at its original URL rather than losing the post.
Publish the media first, then post it
The recommended flow is to put a recording on catcrafts.net before writing the
post, and use that URL as the post's link. Run tools/publish-media.sh recording.mp4; it transcodes to AV1 and an H.264 sibling, uploads both to the
media mount under the AV1's content hash, and prints a
https://catcrafts.net/media/<hash>.h264.mp4 URL to paste into the post. The
H.264 one on purpose: a post's link is fetched raw — Lemmy apps and browsers play
that exact file, with no negotiation in front of it — so it must be the encoding
everything can play. (An AV1 link posted before this existed drew "bad media
error" reports from iPhones within hours.)
fetch-media.sh recognises its own origin and adopts such a URL: it rewrites
it to /media/<hash> and probes the local file for dimensions, downloading
nothing. That is not just an optimisation — it removes the whole class of build
failure where the media step depends on a third party. A 167 MB recording on a
file host once hit MAX_BYTES (64 MB), so the entry kept its original URL, and
the deploy then failed the e2e "media origin" check on a file that was sitting on
our own disk the entire time.
Adopting a <hash>.h264.mp4 URL swaps the AV1 sibling back in as the page's
primary <source> when it is on the mount, keeping the H.264 as the fallback
<source>. So the post links the compatible file, while browsers that can take
AV1 download the small one — the codecs parameter on the first source is what
lets the rest skip it.
The transcode matters as much as the hosting. Phone recordings are wildly
oversized for what they show — that same 167 MB clip was 78 s of a dark room at
17 Mbps, and denoising into AV1 gives the same picture in 15 MB. It also bakes
in the rotation: phones record landscape and attach a display matrix, so an
untouched file reports 1920x1080 while playing portrait, and the width/height
attributes then reserve exactly the wrong box. (fetch-media.sh swaps the
dimensions when it sees a quarter-turn matrix, so a straight-from-phone mirror is
correct too — but transcoding means nothing downstream has to know.)
Two things to know about publishing AV1:
- Browsers without AV1 (Safari before 17, Apple hardware older than A17/M3) get
the
<hash>.h264.mp4sibling: on the site via the second<source>, on the fediverse because that sibling is the posted URL.--rawskips the transcode and the fallback both, so a raw AV1 upload recreates the will-not-play problem — use it for files that are already universally playable. - Lemmy's
pict-rswill not generate a thumbnail from an AV1 file, so a self-hosted video usually arrives with noposter.publish-media.shuploads a poster frame beside the video, named<video-hash>.poster.webp, andfetch-media.shfalls back to that sibling when the instance supplied nothing. A thumbnail the instance did provide always wins. (Posting the H.264 URL also means pict-rs can thumbnail it again, so instance thumbnails come back.)
An own-origin URL naming a file that is not on the mount is deliberately left
pointing at its original URL rather than rewritten. That is a post published
without its media, and failing the e2e origin check loudly beats shipping a 404
inside a <video> tag.
fetch-media.sh wants ffprobe (Arch: ffmpeg) to read pixel dimensions,
which become the width/height attributes that stop the page reflowing as
several 5 MB recordings arrive. It degrades to no dimensions without it —
so tools/e2e.sh asserts they are present, and a build host missing the package
fails at the e2e gate instead of quietly shipping a janky page.
Only the post's main media is used. Lemmy distinguishes the media a post is
(post.url, with post.thumbnail_url as its generated still) from images merely
embedded in the body. Only the former is mirrored — a post whose pictures are
body-only renders as text, which is correct. Videos get the thumbnail as their
poster; without it a preload="metadata" video is a black box until someone
presses play.
Thread links point at the community's instance, not the author's. A post's
ap_id is on the account's own instance, but the discussion is in the community,
which is federated elsewhere and has its own local id for the same thread. So
fetch-posts.sh resolves each ap_id through the community instance's
resolve_object and stores that URL. One request per post, at build time; a
failure falls back to the ap_id, which still reaches a readable copy.
Media is mirrored, not hotlinked. Three reasons: the privacy notice says everything the browser loads comes from catcrafts.net and should stay true; hotlinking would send every visitor's IP to the source instance; and these posts are their media, so a deleted upstream file would gut the page. Filenames are the content hash, which is why the cache lifetime can be a year.
Payments: Mollie setup
The rail is Mollie (bunq.me was measured and disqualified: €500/transaction on cards and no method at all for a non-EU buyer at phone prices — it is a P2P tool; the bunq client remains in the tree, unused, in case an account sweep is ever wanted). The server needs exactly one secret: the Mollie API key.
# Keys live in the Mollie dashboard: Developers -> API keys. A test_… key
# works against the real API from the moment the account exists — verify the
# whole flow with it BEFORE swapping in the live_… key.
install -d -m 0755 /etc/catcrafts
cat > /etc/catcrafts/payments.env <<'ENV'
MOLLIE_API_KEY=test_your-key-here
ENV
chmod 0600 /etc/catcrafts/payments.env
systemctl restart catcrafts-server
journalctl -u catcrafts-server | tail # should say "payments: mollie"
Without the env file the server starts with payments off: the whole site works, the product page renders, and checkout answers 503 with an honest message — degraded, not down.
Mechanics worth knowing:
- The reconciler polls each open order (
GET /v2/payments/{id}) every 10 s while fresh, backing off with age.?redirectback from Mollie is ignored by design — only the authenticated poll moves an order to paid, and the paid event records the method (ideal,creditcard, …) in the ledger. - Mollie payments EXPIRE. A payment that reaches canceled/expired/failed lapses the order automatically — the buyer just orders again.
- Card money stays disputable for months even after "paid": before shipping a
large or exported order, glance at the
viacolumn in--orders. iDEAL and bank transfers are final;creditcardis the one with a tail. - Mollie onboarding reviews the shop: the imprint (KVK, contact address), terms and privacy pages must be real before they approve live payments. They are — and e2e now fails the build if a PLACEHOLDER marker ever reaches a rendered page again.
Shipping rates: Sendcloud (optional)
Without configuration, shipping is priced by the three-zone table in the compiled-in product data (NL / EU / world, Catcrafts.Shared-Content.cppm) — honest flat rates you set. With a Sendcloud account, the server fetches the real per-country prices of one shipping method daily and uses those instead, falling back to the zones for any country the method does not cover:
# credentials from Sendcloud: Settings -> Integrations -> API
cat >> /etc/catcrafts/bunq.env <<'ENV'
SENDCLOUD_PUBLIC_KEY=...
SENDCLOUD_SECRET_KEY=...
SENDCLOUD_METHOD='PostNL Parcels non-EU,DPD Home'
ENV
systemctl restart catcrafts-server
journalctl -u catcrafts-server | grep shipping: # "table refreshed (N countries...)"
SENDCLOUD_METHOD is a comma-separated list of name substrings, merged in
order with the FIRST match per country winning — put the postal method
first so non-EU destinations get post rates (a courier method that also
covers Norway or Switzerland would otherwise price them at courier rates,
€54 instead of €19), and the courier second to fill the EU. The fetched table is
cached next to the orders file so a restart during a Sendcloud outage keeps
the last known prices. Like the bunq client, this integration is UNTESTED
against the live API until credentials exist — the response parser is covered
by --selftest, the fetch around it is thin.
The buyer sees whatever the server will charge: the checkout page embeds the active table into its live total, and the amount is computed server-side at order time from the same data.
Invoice signing (GPG)
Paid orders offer a clearsigned markdown invoice at /order/<token>/invoice.md.
Numbering continues the pre-shop administration: one series per customer — a
random UUID as the customer number, invoices counting sequentially within it
(f57c6512-…-3), keyed by the buyer's email. Art. 226(2) permits "one or more
series"; completeness is provable by reconciling the append-only ledger against
the payment provider's records. The signature makes the invoice verifiable
forever, independent of this server — which is why the order page tells buyers
to download it rather than promising to host receipts indefinitely.
One-time key setup on the server, as the service user:
sudo -u catcrafts env GNUPGHOME=/var/lib/catcrafts/gnupg \
gpg --batch --passphrase '' --quick-gen-key 'Catcrafts invoices <invoices@catcrafts.net>' default default never
# export the PUBLIC key and commit it to the repo so buyers can verify:
sudo -u catcrafts env GNUPGHOME=/var/lib/catcrafts/gnupg \
gpg --armor --export invoices@catcrafts.net > invoice-key.asc
Then in /etc/catcrafts/payments.env:
INVOICE_GPG_KEY=invoices@catcrafts.net
and Environment=GNUPGHOME=/var/lib/catcrafts/gnupg in the service unit (see
catcrafts-server.service). The key has no passphrase because the service signs
unattended; the keyring lives in the 0700 StateDirectory. With a key configured,
a signing failure is a 500 — an unsigned invoice is never served by accident.
Without one (dev), invoices carry a visible UNSIGNED marker.
Reading the orders ledger
catcrafts-server --orders /var/lib/catcrafts/orders.jsonl
orders: 2
reference status total cc created token
CC-3F9A2C paid 595.00 NL 2026-08-04T14:02:11Z 3f9a2c…
CC-91B04D awaiting_payment 534.34 CA 2026-08-04T15:40:03Z 91b04d…
Manual transitions exist for the cases automation cannot see — a payment bunq confirmed out-of-band, the parcel handed to the carrier, a refund:
catcrafts-server --orders /var/lib/catcrafts/orders.jsonl --mark-paid <token>
catcrafts-server --orders /var/lib/catcrafts/orders.jsonl --mark-shipped <token>
catcrafts-server --orders /var/lib/catcrafts/orders.jsonl --cancel <token>
Each appends a status event to the log — nothing is ever rewritten, so the file remains its own audit trail. Deleting personal data on request is an edit of the fields the seven-year fiscal retention does not cover.
Verifying a deploy
systemctl status catcrafts-server
curl -s localhost:8081/api/healthz # ok + content counts
# real status codes, which a client-side router cannot produce
curl -o /dev/null -w '%{http_code}\n' https://catcrafts.net/nope # 404
curl -o /dev/null -w '%{http_code}\n' https://catcrafts.net/blog # 301
# the SEO check: content present with no JavaScript involved
curl -s https://catcrafts.net/projects | grep -c '<script' # 0
curl -s https://catcrafts.net/projects | grep -o '<title>[^<]*'
Analytics
Server-side only — the privacy policy promises request logging and nothing else, so there is no client-side analytics anywhere on the site. GoAccess (Debian package) turns Caddy's JSON access logs into two static HTML reports, each with its own persistent DB and ingest ledger:
https://catcrafts.net/analytics/— public, censored. Visitor IPs are anonymized at ingest (last octet zeroed before anything reaches its DB), no HOSTS or full-URL REFERRERS panels, and log lines matchingCENSOR_REin the script never enter its DB at all — the public tier cannot leak what it never ingested. ExtendCENSOR_REwhen the shop launches so order/payment URLs can never surface; keep secrets out of URL paths regardless (query strings are already stripped).https://catcrafts.net/analytics/private/— uncensored (basic auth, hash in the Caddyfile): full IPs, all panels.
Raw logs keep full IPs either way — that is the request logging the privacy policy declares; per-IP forensics work from the logs and the private tier, never from the public page.
Three pieces, all in deploy/:
catcrafts-analytics→/usr/local/bin/— ingests each rotatedcatcrafts.net-*.log.gzexactly once into a persistent GoAccess DB (/var/lib/goaccess/db, tracked in/var/lib/goaccess/ingested), then renders the report from DB + live log. The live file is never persisted, so its lines don't double-count when Caddy rotates it. History therefore survives log deletion: the DB keeps aggregates forever.catcrafts-analytics.service— oneshot, runs ascaddy(owner of the 0600 logs).catcrafts-analytics.timer— hourly at :07.
Bot filtering is the load-bearing part: measured on real traffic, 57% of
requests were headerless vulnerability scanners and another 19% self-declared
bots (mostly ClaudeBot) — only ~24% human. --ignore-crawlers --unknowns-as-crawlers drops both groups. The flags in the script apply at
ingest time and the DB stores aggregated data, so changing filters later only
affects new lines — re-ingesting history means deleting
/var/lib/goaccess/{db,ingested} and letting the next run rebuild from
whatever raw logs retention still holds (a year, per the Caddyfile).
apt install goaccess
install -m 755 deploy/catcrafts-analytics /usr/local/bin/
install -m 644 deploy/catcrafts-analytics.{service,timer} /etc/systemd/system/
mkdir -p /etc/goaccess /var/lib/goaccess /var/www/analytics /var/www/analytics-private
install -m 644 deploy/goaccess-browsers.list /etc/goaccess/browsers.list
# own IPs to keep out of the numbers - host-only file, NOT in this repo
echo "203.0.113.7" > /etc/goaccess/exclude-ips
chown -R caddy:caddy /var/lib/goaccess /var/www/analytics /var/www/analytics-private
systemctl daemon-reload && systemctl enable --now catcrafts-analytics.timer
Running it locally
tools/dev.sh # build both products and serve on :8080
tools/dev.sh --no-build # reuse what is already in bin/
That is the whole site: Caddy in front, the backend behind, static assets from
disk, cross-origin headers scoped exactly as in production. Ctrl-C stops both.
Orders from the session go to a temp file and are discarded on exit; the
fake payment rail is active — touch <workdir>/orders.jsonl.fake-paid plays
the part of the customer paying.
Do not run catcrafts-server --serve alone and expect a working site. It
serves pages only — static assets are Caddy's job — so /styles.css 404s and
every page renders unstyled. That looks broken but isn't; it is the deployment
split working as designed.
Other useful commands
tools/fetch-posts.sh # pull posts from the allowed communities
tools/fetch-media.sh [dir] # mirror their media locally (run after the above)
tools/fetch-rates.sh # ECB reference rates for the indicative prices
tools/e2e.sh # 143 HTTP checks against a real server; the CI gate
crafter-build --local -r # the wasm app alone, no backend, on :8080
<server>/catcrafts-server --selftest # ~140 in-process assertions
<server>/catcrafts-server --routes # status + title for every route
<server>/catcrafts-server --render /projects # dump one page's HTML
<server>/catcrafts-server --orders FILE # the orders ledger + manual transitions
If a bin/ glob matches two directories
The variant directory name embeds a config hash, and --local and
non---local builds hash differently: --local resolves the Crafter libraries
from sibling working trees, a plain build fetches them from Forgejo. They are
genuinely different configurations and both land in bin/.
Mixing them leaves two directories, and any script globbing for one picks
arbitrarily — which in practice means testing a stale binary and believing the
result. Every script here refuses to guess and tells you to rm -rf bin. Pick
one mode and stay in it.