# rustnim — a Rust → Nim transpiler ## Status **Milestone 1 is reached: all of `base16ct` goes through.** Every one of its source files transpiles byte-for-byte as published, `alloc` half included, and its decode and encode output is byte-identical to rustc's. 33 differential cases, 29 behavioural and 4 rejections, plus 6 unit/integration tests. All green. Run `cargo test`. Passing today: functions, `impl` methods, trait impls (formatting traits and `From`), structs, enums (C-like and data-carrying), `Option`/`Result` with `?`, closures, `unsafe`, slice iterators (`iter`/`iter_mut`/`enumerate`/`zip`/`chunks_exact`/ `chunks_exact_mut`/`windows`), borrowed slices as values and return types, `let`/`let mut`, the full integer and float operator set at exact widths, `as` casts, `if`/`while`/`loop`/`for`, `match` including patterns that bind, `Vec`/slices/arrays, type aliases (including generic ones), function-typed parameters (`impl Fn(A) -> B`), multi-file input, `#[cfg]` evaluation, and `println!`/`format!` with `{}`, `{:?}`, `{:x}`, `{:b}`, positional and inline-named arguments, and zero/space padding. ## Why this exists We tried [tarekwasfy01/Code-Transpiler](https://github.com/tarekwasfy01/Code-Transpiler), which advertises `rust` as a source language, on the `base16ct` crate. It emits empty files and exits 0. The full investigation is in [`findings/`](findings/) and is published at https://rickub.com/nandi/code-transpiler-rust-frontend-findings The decisive finding, and the reason this is a new project rather than a patch: its Universal AST cannot represent Rust. `defaultSemanticTypeContract()` in `internal/backend/semantic_program.go:85` is hardcoded to ``` numeric: binary64, integer_width: unknown, truth: r_compatible, ownership: unknown, index_base: 1 ``` and `semantic_document.go:1014` *validates* that every contract equals exactly that, while `typed_operation.go:46` rejects any value model that is not `tagged_dynamic_binary64`. There is no integer width and no ownership in the model at all. Code like `base16ct`'s constant-time decoder — ```rust ret += (((0x2fi16 - byte) & (byte - 0x3a)) >> 8) & (byte - 47); ``` — depends on exact 16-bit signed wrapping and arithmetic shift. Lowering that into a 1-indexed dynamic float64 model produces silently wrong answers. So the first rule of this project is the one that codebase broke: > **Never approximate a semantic you cannot represent. Fail loudly instead.** `src/ty.rs` already does this: `i128`/`u128` are rejected with a reason rather than widened or truncated. ## Architecture ``` Rust source ──syn──> syn AST ──lower──> Nim source ──nim c──> binary ``` **The frontend is `syn`, deliberately.** Hand-rolling a Rust grammar is how the other project went wrong; a correct parser is not the interesting part of this problem. The interesting part is the lowering, which is where all the work goes. Planned modules: | file | role | state | |---|---|---| | `src/ty.rs` | Rust type → Nim type, exact widths, explicit rejections | written | | `src/lower.rs` | items, statements, expressions → Nim | written | | `src/fmt.rs` | `println!`/`format!` format-string handling | written | | `src/prelude.nim` | `Option`/`Result`/panic/`Display`/`Debug` runtime | written | | `src/main.rs` | CLI: `rustnim -o ` | written | | `tests/differential.rs` | the runner described below | written | ### Enums, `Option` and `Result` A C-like enum becomes a plain Nim `enum`, which compares, orders and `case`-checks the way Rust's does. A data-carrying enum becomes a Nim object variant — a discriminant enum plus one branch per variant — which is the same shape the prelude already uses for `Option` and `Result`. Nim requires the branches of a variant object to have distinct field names, so each payload field is prefixed with its variant. `match` takes one of two forms. Arms that neither bind nor destructure become a Nim `case`, which is exhaustiveness-checked the way Rust's is. Arms that do bind become an `if`/`elif` chain with the bindings emitted as `let`s, because Nim's `case` cannot destructure. The chain always ends in an arm that panics: Rust proved it unreachable, but Nim cannot see that, and leaving the chain open would silently fall through instead. `Ok`, `Err` and `Some` are emitted with their full type arguments (`rsOk[T, E](v)`), because Nim cannot infer `E` from an `Ok(v)` alone. That is why the expected type has to reach a `match` arm as well as a `let`. `?` expands to statements — a temporary, a discriminant test, and an early `return` — which are emitted ahead of the line being built. Rust inserts a `From::from` on the error there; we accept only the case where the two error types already agree, rather than assume a conversion is the identity. `?` in a `while` condition is rejected: the early return would run once before the loop rather than on each iteration. ### Trait impls A `Display` impl becomes `proc rsDisplay(self: T): string`. Rust's `Formatter` is a sink and the observable result of `{}` is exactly the bytes written into it, so a write through the formatter **appends** to that string — a `fmt` body may write repeatedly, and `UpperHex` writes once per byte in a loop. A body that does anything else with the formatter — padding, precision, `debug_struct` — is rejected, because those change the output and this model does not carry them. `Debug`, `LowerHex`, `UpperHex`, `Binary` and `Octal` work the same way. Writing into a string cannot fail, so `?` on a formatter write is a no-op. `?` on anything else inside a `fmt` body *can* fail, and `format!` panics when a formatting impl returns an error — so that is what the error branch does, with std's own message. `{:x}` on an integer formats its two's-complement bit pattern; on any other type it calls that type's own `LowerHex` impl. Those are different operations, so a radix format on an argument of unknown type is rejected rather than guessed. `impl From for B` becomes a conversion proc that `.into()` resolves through. A marker trait with no items generates nothing: we do not model trait resolution anywhere, so there is nothing for it to affect; a use that actually needed the trait (a `dyn`, a bound) is rejected where it appears. Any other trait impl is rejected. Methods are keyed by `(receiver type, name)`, not by name alone — two types may define the same method, and Nim tells them apart by overload resolution on the first parameter. `fmt::Error` is *not* the same type as a crate's own `Error`. Collapsing a qualified path to its last segment merged them, which was a real soundness bug; `core::fmt`'s types are now recognised by their qualified name. ### Slice iterators are resolved to one index loop Rust's slice iterators are lazy and compose. Nim's `for` is over one sequence, so a chain of adaptors is resolved into a small IR and emitted as a single index loop in which **each binding is an lvalue into the original container**. That is what makes `*d = v` through `iter_mut()` write back to the caller's slice instead of to a copy, and what lets `chunks_exact(2)` hand out a window that indexes straight into the source with an offset. Only adaptors with an exact index-loop equivalent are accepted. `map`, `filter` and `take_while` are rejected rather than partially honoured: silently dropping an adaptor would change which elements the loop visits. `zip` stops at the shorter side, as Rust's does — that is a test, not an assumption (`tests/cases/023`). ### Borrowed slices are views, not copies `&[T]` is a borrow. Nim's experimental view types model exactly that, including returning one from a proc: writing through the returned view is visible in the original buffer. That was probed against Nim 2.2.4 before being relied on, because copying into a `seq` would print the right bytes while silently changing aliasing. Nim does allow a view inside an object and inside an object *field* — both probed, both preserving aliasing — so `Result<&[u8], E>` and `HexDisplay<'a>(&'a [u8])` both work. (An earlier version of this document claimed otherwise; that was wrong.) Two real constraints remain. Nim will not let a `let` borrow out of a local, so `.unwrap()`/`.expect()` on a `Result` holding a view is expanded inline and the binding becomes an alias — a view is a reference, so there is nothing to materialise, and the substituted expression is a plain field access that re-evaluates nothing. And `s.get(a..b)` is an `Option` of a view whose *validity* is what matters: the view and its condition travel together through `ok_or` until a `?` or `unwrap` resolves them into a bounds check plus a binding. Keeping such an `Option` in a variable is rejected with a message saying so. A `let` binding a borrow keeps the view rather than copying into a `seq`: `let res = encode(..)?` names the caller's buffer, and copying would print the right bytes while silently breaking the aliasing. ### Closures and `unsafe` `unsafe` is a permission marker, not a semantic change: it does not alter what the enclosed operations mean. So the block is transparent, and every operation inside still goes through the ordinary lowering and is still rejected if it has no faithful mapping. `unsafe fn` lowers like any other proc. A closure becomes a Nim anonymous proc. Nim's closures capture by reference, as Rust's non-`move` closures do; a `move` closure captures by value, which is a different thing, so it is rejected rather than lowered to the same construct. `impl Fn(A) -> B` is left at Nim's default calling convention, which accepts both a plain top-level proc and a capturing closure — as Rust's `impl Fn` does. `.map`/`.and_then` over an `Option`/`Result` are expanded inline with the closure's parameter aliased to the payload, rather than handed to a generic proc. That keeps the whole thing an expression and keeps a view a view. `&str` is a borrowed view of someone else's bytes, so it maps to `openArray[char]`, not to an owned `string`. Nim accepts a `string` argument for an `openArray[char]` parameter, so a literal still passes straight through. `from_utf8_unchecked` reinterprets a byte view as a character view over the same memory — no copy, no validation, and writes through the original are visible, as in Rust. ### Modules Rust keeps `lower::decode` and `mixed::decode` apart by module; flattening into one Nim module would merge them — they are *different functions*. So the first input is the crate root and each later one is a module named by its file stem, items are emitted as `_`, and a call resolves through an explicit qualifier, then the current module, then what `use` brought into scope, then the root. ### Declaration order Rust has no declaration-before-use rule and Nim does, so every proc is forward-declared between the type definitions and the bodies. Reordering the input instead would not handle mutual recursion. ### Type propagation is load-bearing Rust infers an unsuffixed integer literal's type from context and falls back to `i32`; Nim falls back to 64-bit `int`. So `lower.rs` threads an *expected type* down through every expression — into `let` annotations, call arguments, `match` patterns, compound assignments and both operands of a binary — and annotates every binding it emits. Without that, `let x: u8 = 200; x + 100` means two different things in the two languages. With it, a width the lowering gets wrong becomes a Nim compile error (a loud failure, reported by the runner) rather than a wrong answer. ## Mapping decisions made so far - **Integers**: exact width. `i32`→`int32`, `usize`→`uint`, etc. `i128`/`u128` rejected. - **Indexing**: both 0-based. Direct. - **`&T`** → plain value. **`&mut T`** → `var T` parameter. - **`&[T]`** → `openArray[T]` in parameter position, `seq[T]` when owned. `Nim::owned()` performs that conversion. - **Ownership/borrowck**: ignored. Nim is GC'd; for safe Rust this is sound. - **`Option`/`Result`** → object variants in the prelude. - **`match`** → Nim `case` where the arms are simple, `if`/`elif` when arms have guards or bindings. - **Rust's expression-orientation** maps well: Nim `if`/`case` are expressions too, and a proc's trailing expression is its return value. ### Settled empirically (Nim 2.2.4 vs rustc 1.98.1, both run) 1. **Nim's `shr` on a signed integer is arithmetic**, matching Rust. `int16(-256) shr 8` = `-1` in Nim; `(-256i16) >> 8` = `-1` in Rust. `base16ct`'s decoder depends on this, so it maps directly with no helper. 2. **Nim's fixed-width unsigned arithmetic wraps silently**, matching Rust's `wrapping_*`. `uint8(200) + 100` = `44` in Nim; `200u8.wrapping_add(100)` = `44` in Rust. So `wrapping_add` on an unsigned type is just `+`. 3. **We model rustc's debug profile.** Rust debug builds panic on signed integer overflow; Nim's default build raises `OverflowDefect` on it. Those are the matching pair, so the runner invokes `rustc` without `-O` and `nim c` with its defaults, and `tests/cases/016` pins the behaviour. A Rust panic exits 101 where a Nim Defect exits 1, so every generated module ends with a handler that maps one to the other — otherwise the runner's exit-status comparison would be vacuous. `wrapping_*` is therefore an explicit operation on both sides: unsigned maps to the bare operator (item 2), signed is routed through the unsigned view of the same width. 4. **`char` round-trips.** Rust `char` → Nim `Rune`, confirmed for ASCII and non-ASCII scalars in both `{}` and `{:?}`, and across `as u32` (`tests/cases/014`). ### Still open 5. `checked_*` and `saturating_*` are not mapped yet; they are currently rejected as unsupported methods rather than approximated. 6. Generics (type and const parameters), `move` closures, closure bodies with statements, and trait impls other than the formatting traits and `From` are rejected with a reason. Lifetime parameters are *not* a rejection: they carry no runtime meaning and Nim is GC'd, so `fn encode<'a>(..)` lowers fine. 7. Float formatting matches Rust for ordinary values and for `inf`/`NaN`, but the exponent-form thresholds have only been checked at `1e21`. 8. Functions are scoped by module now, but *types* are still global: two modules declaring the same type name would collide. Relatedly, a crate's own `type Result` is told apart from the builtin `Result` by arity, which is not how Rust resolves it. 9. `String::from_utf8_unchecked` copies, because Nim's `string` is an owned value. Rust's consumes the `Vec` without copying. Observably the same from the caller, but it is a copy where Rust has none. ## Testing: differential, not golden The bar is **behavioural equivalence with rustc**, not that the output looks plausible. For each case in `tests/cases/`: ``` rustc case.rs && ./case > expected rustnim case.rs -o case.nim && nim c -r case.nim > actual diff expected actual ``` A case only counts as passing when both binaries build *and* produce identical stdout *and* exit with the same status. `tests/differential.rs` implements this, and checks each stage separately so a failure says where it went wrong: `rustnim`, `rustc`, `nim`, or `diff`. Three guards exist specifically because of how the other transpiler failed: - `rustnim` exiting 0 while writing **no output file** is a failure. - `rustnim` exiting 0 while writing an **empty output file** is a failure. - An **empty corpus** is a failure, so the runner cannot pass by finding nothing to do. All three have been verified by deliberately breaking the transpiler and confirming the runner goes red. Cases carry directives in leading `//@` comments: | directive | meaning | |---|---| | `//@ reject: ` | `rustnim` must *fail*, with this in its message | | `//@ skip: ` | not run; reported as skipped | | `//@ args: ` | passed to both binaries | | `//@ stdin: ` | fed to both binaries | `reject` cases are how the "fail loudly" rule is tested rather than merely stated: `900`–`904` pin the rejections of `i128`, an unmapped standard-library method, a float→int cast, an unimplemented format spec, and a closure. Run one case with `RUSTNIM_CASE=005 cargo test --test differential -- --nocapture`. Nim is found at `.nim-toolchain/bin/nim` in the repository root or any parent, or via `RUSTNIM_NIM`. ## Toolchain - `rustc` / `cargo` 1.98.1 — system. - Nim 2.2.4 — vendored at `.nim-toolchain/` (gitignored; downloaded from nim-lang.org, not installed system-wide). Binary: `.nim-toolchain/bin/nim`. ## Milestone 1 Transpile `base16ct` 1.0.0 — the crate the other transpiler failed on — and have its decoder produce byte-identical output to the Rust original. **Reached.** `tests/cases/026-base16ct-crate/` transpiles **every source file of base16ct 1.0.0** — `error.rs`, `lower.rs`, `upper.rs`, `mixed.rs` and `display.rs`, each byte-for-byte as published on crates.io, verified with `cmp` rather than by eye — together with `lib.rs`'s `decoded_len`, `encoded_len` and `decode_inner` verbatim. The `alloc` half is on, via `--cfg feature=alloc`. Output is byte-identical to rustc's: ``` lower ok abcd1234 len=4 decode: lower, upper, mixed upper-rej err InvalidEncoding ... upper correctly rejects lowercase oddlen err InvalidLength / invalid Base16 length <- Debug and Display encode ok 6162636431323334 len=8 encode, both cases encode_str ok abcd1234 len=8 closure over unsafe, borrowed &str Ok([171, 205, 18, 52]) decode_vec \ abcd1234 encode_string > the alloc half ABCD1234 abcd1234 HexDisplay {:X} {:x} ``` Everything lowers as written: `dst.get_mut(..decoded_len(src)?)`, `src.chunks_exact(2).zip(dst.iter_mut())`, `*dst = byte as u8`, the returned `&'a [u8]` view into the caller's buffer, `encode(src, dst).map(|r| unsafe { core::str::from_utf8_unchecked(r) })`, and `HexDisplay`'s `UpperHex` impl writing once per byte into the formatter. This is the crate whose six files the transpiler in `findings/` emitted empty output for, while exiting 0.