# RustyRPN RustyRPN is a project to create and maintain a business management application for a small Swedish Aktiebolag. ## The business The company in question is *Rent & Petroleum Nordic AB* with a yearly turnover of about 20 million SEK. RPNAB has two main revenue generators: 1. Selling gasoline and diesel at the local airport both to retail, as well as to the car rental firms located at the airport. 2. Acting as a franchisee to Enterprise Rent-A-Car. RPNAB is historically a family company and is at the moment driven by two cousins; Johan (CEO) and Jakob (chairman of the board) Rönnbäck. For a few years RPNAB had a daughter company called Recamp Nordic AB that everything related to car rentals was delegated to, but it was recently absorbed back into the main company. ## The maintainer - Official position within the company is board-of-director as well as owner of 60% of the private equity - Have some minor experience with development (mainly node.js) - Is using this project as an opportunity to learn Rust, as well as AI-assisted development - Prefer a TDD approach and functional code style (used ramda library while writing node.js code) - Uses OmniFocus and a GTD inspired workflow for keeping track of tasks ## The application Code will eventually be kept on a private gitea instance (for issue handling), but everything is to be considered open source. No expectations of assistance with writing code, but if the project and the company is successful a dream of the maintainer is for other companies and developers to make use of it. However, due to the bespoke nature of the functionality provided this seems unlikely. ## Toolchain - Rust via rustup, stable channel (installed with `brew install rustup-init; rustup default stable` on macOS; rustup's shims live in `/opt/homebrew/opt/rustup/bin` — keep it on PATH). - Build and test from the `Application/` directory: `cargo build`, `cargo test`, `cargo fmt --check`, `cargo clippy`. - Python 3 (stdlib + openpyxl + pypdf) is used only by `Application/scripts/` for regenerating sanitized test fixtures; it is not part of the product toolchain. ## Technology The whole stack is Rust: one language, one toolchain, one test runner, from database to browser. Motivation for choosing Rust is maintainability, not speed or efficiency. - Rust: stable channel, edition 2024. Minimum supported version: 1.85 (the first release with edition 2024). Crate versions are decided at `cargo add` time and pinned in Cargo.lock; this document records choices, not versions. - Async runtime: tokio. - CLI: clap (derive API). - Web server & JSON API: axum. - SPA: Leptos, served by the same binary as the API. - Database: MariaDB (protocol-compatible MySQL). Client: sqlx, with mysql_async as fallback. The DB server itself is out of scope; given connection details in config, the application creates its own tables from embedded migrations (schema_migrations table). The v1 domain schema (files, customers, batches, cards, transactions, invoices, invoice_items) is documented in Documentation/schema.sql; it is the initial embedded migration. - Config: TOML file (see config.template.toml), loaded via --config flag or $RUSTYRPN_CONFIG. Real config files are gitignored. Daemon state (pidfile, logs) and database backups default to ~/.config/rpn (backups in its backups/ subdirectory); both are overridable in config. - Serialization: serde / serde_json. - Logging: tracing (level via RUST_LOG). - CSV/TSV files: csv crate; non-standard raw-export dates are parsed by our own module, test-covered. - Dates: chrono. - Errors: thiserror for domain errors, anyhow in binaries. Repository layout (Cargo workspace, root: `Application/`): - src/core – domain logic: ingest, invoices, Fortnox, car registry - src/cli – command line interface - src/server – axum server: JSON API + SPA assets - src/web – Leptos SPA Build & test: - cargo build / cargo test at the workspace root - Tests needing a database use config.test.toml and a scratch schema; tests must not assume pre-existing data - The binaries are built on FreeBSD (pkg rust); they run inside a jail Hosting: - Caddy reverse-proxies the web interface (TLS) from a separate jail (running caddy is out of scope) - The CLI is used from the LAN only (e.g. via SSH); nothing is exposed for it. Authentication & Authorization: - Portal authentication (v1): login form with email + password (argon2id) - Server-side session in MariaDB behind a HttpOnly/Secure/SameSite=Lax cookie, rate-limited, no user enumeration - Roles: - director (Jakob & Johan, full access) - employee (access to appropriate internal functions) - customer (own company's invoices only — every query scoped to the session's company) - OIDC and passkeys are future options, not part of v1. ## Data formats Three document types flow into the application. Samples for all of them live under `Application/data/test_input/` in one directory per source. Samples contain real customer and card data: they are gitignored, must never be committed, and their values must never be embedded in committed test files. Tests may *reference* sample files by path (the files stay local), or use synthetic/sanitized values in the repo. Sanitized copies of the samples — committed, safe for any repo consumer — live under `Application/data/fixtures/` (epsilon batch files 405-412 with synthetic card/customer ids, batches renumbered 9405+ and dates +2y; a cumulative slice, batches 5001+, dates -6y; tsdrms xlsx files with synthetic R/A / DBR / location ids, dates +2y; subfranchise statements as extracted text with all identity and amounts replaced). Amounts, GL codes and descriptions are kept verbatim where they carry no personal data; cross-value arithmetic in the subfranchise fixtures (EUR × rate = SEK) is deliberately NOT preserved. Regenerate with `python3 Application/scripts/sanitize_samples.py` (reads the local raw samples) and verify sanitization with `check_fixture_leaks.py` (exit 0 = clean). Feature branches should build parser tests against the fixtures. The raw formats are canonical: the application parses what the source systems deliver. Filenames carry metadata (batch / month / invoice number) and are the natural keys for idempotent re-ingestion — ingesting the same file twice must not duplicate data. Date and number formats differ per source; never assume: | Source | Date format | Example | Number format | |---|---|---|---| | epsilon | US style, no zero padding: `M/d/yyyy h:mm:ss AM/PM` | `3/16/2026 6:08:43 AM` | `561.24` | | tsdrms | EU style: `DD/MM/YYYY` | `07/01/2026` | `-540.54` | | subfranchise PDF | Month name + year in text | `January 2026` | European: `1.387.732,10`, `(35.070,15)` for negatives | ### 1. Epsilon fuel station exports Sales transactions from the station's cash register system Epsilon - Filename: `raw-export-from-epsilon-.txt`, e.g. `raw-export-from-epsilon-405.txt`. One file = one batch; the filename number equals the `Batch number` field of every row in the file. - Batches are numbered sequentially (two per month in samples) and span several days each (e.g. batch 409: 2/1–2/9). Cumulative exports can contain many batches (sample: batches 1–406, ~138k rows, 2019→2025). - File format: ASCII, tab-separated, every field double-quoted, CRLF line endings, header row, trailing newline. 16 fields: | Field | Example | Notes | |---|---|---| | Date | `3/16/2026 6:08:43 AM` | US format, see table above | | Batch number | `405` | matches filename | | Amount | `561.24` | SEK; = Volume × Price | | Volume | `31.18` | liters | | Price | `18.00` | SEK/liter | | Quality | `1001` | code: 1001 = unleaded, 4 = Diesel, 0 = zero-value record | | QualityName | `95 Oktan` | empty for Quality 0 | | Card number | `549543******5778` | consumer cards are masked; contract cards appear **unmasked** and are exactly the rows with a non-empty Customer number — personal data, store as delivered | | Card type | `549543******5778` | equals Card number in all samples | | Customer number | `1861` | contract customer id; empty for retail | | Station | `97254` | single station in samples | | Terminal / Pump | `1` / `2` | small integers | | Receipt | `004109` | zero-padded 6 digits | | Card report group number | `4` | | | Control number | `126301` | alphanumeric; empty on some rows | Known edge cases (from samples — verify against new real files before changing the parser): - Quality 0 rows are zero-value records (amount, volume, price all `0.00`, empty name) - canceled fuelings and testing events - No negative amounts in any sample. ### 2. TSDRMS exports (Enterprise rental system) - Filename: `YYYY-MM.xlsx` (monthly) or `YYYY-MM-DD - YYYY-MM-DD.xlsx` (period), e.g. `2026-01.xlsx`, `2026-01-01 - 2026-07-31.xlsx`. One sheet, named `Sheet`. - Content: a **general ledger journal export** (double-entry), *not* a list of rentals. Each rental appears as paired debit/credit rows. 11 fields: | Field | Example | Notes | |---|---|---| | GL Account # | `2050` | `2050` DEPOSITS RECEIVED, `8000` MASTERCARD, `1030` WIRE TRANSFER PAYMENT, `1511` ACCOUNTS RECEIVABLE, `2611` TAX TYPE 1, … full code list unknown — derive from data | | Description | `DEPOSITS RECEIVED` | GL account name | | Location | `LLAT73` | Enterprise location code | | CODE | `M`, `R1`, `P`, `D`, `CDWTPI` | product identifier used to map rows to bookkeeping accounts | | R/A # | `LLAT62-3986i5` | rental agreement id; format varies (also `345670C`) | | Transaction Date | `07/01/2026` | DD/MM/YYYY | | DBR | `LLAC61-1260` | meaning unknown | | Cutoff Date | `07/01/2026` | DD/MM/YYYY | | Debit / Credit | `-1.25` / `1.25` | numeric; stored as IEEE floats in the XML (e.g. `-540.53999999999996`) — round to 2 decimals on ingest | | Product Type | `VEHICLE` | `VEHICLE` / `NO PRODUCT` | - **Partitioning quirk (important):** monthly files are partitioned by *Cutoff Date*, not Transaction Date. All 2797 rows of the January 2026 file have January cutoffs, but 23 of them have December 2025 transaction dates. Never assume transaction month = file month. - Scale: ~2.8k–3.6k rows per month; ~19.9k rows for a 7-month period file. ### 3. Subfranchise statements (PDF) Monthly settlement statements from Enterprise (Shared Mobility Sverige Filial, org.nr 516411-7920) for the sub-franchised business. One PDF per settlement month, received ~1–2 months after the period ends. - Filename: ` -- - Subfranchise Statement - .pdf`, e.g. `2026-03-02 -- 821100844 - Subfranchise Statement - 2026-01.pdf`. Invoice numbers are sequential (821100844, 821100853, …). - 1-page A4, **text-based (not scanned)** — text extraction works. Choose a Rust PDF text-extraction crate when this feature starts; layout stability is verified for six samples (2026-01 → 2026-06) only, so parse defensively and re-verify against each new statement. - Header block: sender name/address/org.nr, settlement period (month name + year), location code (`LLAT`), Invoice #, Date, Partner (`Recamp Nordic AB`), Reference (e.g. `Subfranchise Statement - Luleå January 2026`). - Body: line-item table with columns `DESCRIPTION | EUR | @daily rate | AMOUNT SEK`. Numbers use the European format from the table above; zero cells read `€ - 10,66 -`. - Known line items (order can vary; some months contain zero rows; names can embed the month, e.g. `InMoment SQI - January 2026` — match by pattern, not exact string): All Revenues, Excluded revenues, Gross revenues less exclusions, Royalty Fee (7%), Central Invoicing Fee (2%), Marketing Fee (1%), reservation-fee line, GF Learning Center, InMoment SQI, Cross Border Debit, Bad Debt Reserve (1%), subtotals, Total fees due, VAT (25%), adjustments, Total amount due to Franchisee. ### 4. Voucher mapping reference (Fortnox) `Application/data/other_resources/` (gitignored) holds a copy of the Apple Numbers spreadsheet the maintainer currently uses to turn tsdrms data into bookkeeping vouchers. It is the reference for the voucher feature: read it to understand the mapping; never modify it. It maps tsdrms `Description` values to Swedish voucher descriptions in these groups: sales (`Fordran - Biluthyrning - Faktura` / `- Shared Mobility`, `Bränsleförsäljning`, VAT-liable / VAT-exempt sales), products and fees (`Försäkringsprodukter`, parking and traffic-violation fees, airport fee, road assistance), damage billing (`Skadedebitering`), write-offs (`Nedskrivning`), card-network and Enterprise reconciliations (`Avstämning - American Express` / `- Windcave` / `- Enterprise`), and output VAT (`Utgående moms, 25%`). Output includes `Export ID` and `Bankkonto` per line. The complete mapping lives in the spreadsheet's formulas and is *not* in the repo. When the voucher feature starts, its first task is to port the mapping into versioned data in the repo (e.g. a TOML table) before writing code against it. ## Development - Two target interfaces - Command line: for scripting & automation, and efficiency for advanced users (i.e. myself) - Web based (SPA) for regular employees and customers - Code - Main priority is maintainability and user facing interface stability - Style should emphasize ease of understanding as well as long term stability - Keep source code clean by mainly keeping descriptive documentation in the git commit messages rather than inline - Use generative AI for writing docs, tests, and code - Use a TDD inspired approach by always writing tests before implementation code - Liberal use of checklists to keep both human and AI on track ### Workflow Use a test-driven-development inspired approach, specifically the following workflow: 1. Create a named git branch for a new feature, improvement, or bugfix 2. Create specification for changes to be made 2.1 Discuss the goal; developer and AI-model discuss a new feature, fix, or improvement to be done 2.2 AI-model asks clarifying questions until a clear understanding is shared between human and AI on what the goal is 2.3 Agree on a checklist of items to test against to help keep the AI on track 3. Write a failing test 4. Commit test to git 5. Write code 5.1 Write the minimal code to pass one of the tests 5.2 Run test and verify test passes 5.3 Create a git commit per passing test 5.4 Iterate for each test until they all pass 6. Verify everything on the checklist has been completed 7. Update roadmap (Documentation/roadmap.md) - roadmap.md tracks what is done, in flight, or planned - readme.md is the source of truth for design decisions: any task that changes a decision (stack, format, workflow) updates it in the same commit 8. Provide a short list of suggestions for improvements and include a suggested global next step ### General guidelines for AI - Do not - commit secrets or private data - commit config files (except for template) - commit anything under data/ **except** `data/fixtures/`, which holds sanitized sample data that is committed deliberately (see §Data formats); run `scripts/check_fixture_leaks.py` (exit 0) before ever touching fixtures - add dependencies without asking - refactor existing code as part of an unrelated task - put secrets or real customer data in tests - Git commit messages - Follow best practices regarding format and context, but also keep in mind the commit messages are the main code documentation - Always ask for feedback before actually committing to `main`; on a feature branch, commit per passing test as part of the workflow without waiting for feedback - Be concise ## Roadmap and status of functionality The roadmap lives in Documentation/roadmap.md: per-item status (done / red / spec / idea), dependencies, acceptance criteria, and near-term ordering. Workflow step 7 updates that file; this document remains the source of truth for design decisions (stack, formats, workflow) — a task that changes a decision updates readme.md, a task that changes *what is built next* updates roadmap.md. ## Glossary - aktiebolag: limited company - besiktning: yearly car inspection mandated by Swedish law - Fortnox: suite of accounting and bookkeeping software - tsdrms: rental management software used by Enterprise and its subfranchisees - Epsilon: cash register system used to administer fuel sales