This document is a comprehensive, implementation-oriented reference for Magpie v0.1. It consolidates language grammar/syntax, core semantics, compiler architecture, CLI arguments, and the LLM-first design rationale.
- Design Goals and Core Decisions
- LLM-First Rationale
- Language Overview
- Lexical Grammar
- File Grammar and Canonical Structure
- Declarations
- Type System
- Instruction Model
- Opcode Reference (Value + Void)
- Semantics and Safety Rules
- TCallable: Why It Exists and How It Works
- ARC + Ownership: Why Both
- Compiler Architecture
- Compiler Pipeline
- Artifacts and Emission Kinds
- Diagnostics Model
- CLI Arguments and Commands
- Configuration Resolution Order
- Build/Test/Run Playbooks
- Examples
Magpie v0.1 is designed around a few non-negotiable engineering constraints:
- Deterministic, explicit surface language
- Minimal hidden behavior.
- Explicit control-flow and explicit ownership-related operations.
- Canonical formatting (CSNF) for stable text representation.
- LLM-agent compatibility
- Structured, low-ambiguity syntax for deterministic generation and repair loops.
- Progressive disclosure and compact machine artifacts.
- Token-budget-aware output behavior.
- Static safety + predictable runtime
- Rust-like ownership/borrowing constraints for aliasing/mutation safety.
- ARC-managed heap lifetimes for deterministic reclamation (no tracing GC pause model).
- Pipeline observability
- Explicit staged compiler pipeline.
- Rich diagnostics, graph artifacts, and JSON envelopes for tool/agent integration.
- Separation of concerns
- Parse/resolve/type/HIR/ownership/MPIR/codegen/link are distinct phases with explicit failure surfaces.
This section summarizes the core LLM-first engineering rationale as applied in Magpie tooling and language design.
Claim (paraphrase): Transformer-based systems operate under bounded context and practical token budgets.
Evidence:
- Original Transformer self-attention scaling with sequence length.
- Efficient Transformer survey on scaling constraints.
- LongBench: practical long-context degradation and retrieval/compression usefulness.
Magpie design implications:
- Budgeted machine-readable outputs.
- Small/local code and progressive disclosure.
- Prefer targeted artifacts (
.mpdbg, graphs,.mpir) over full-context dumps.
Claim: Long sequences reduce reliable focus on relevant middle-span details.
Evidence:
- “Lost in the Middle” shows severe position effects for relevant information in long contexts.
Magpie implications:
- Locality constraints, explicit dependencies, compact structured summaries.
- Encourage retrieval of only relevant slices rather than monolithic context.
Claim: Tokenization and boundary behavior can destabilize model output.
Evidence:
- Tokenization robustness degradation results.
- Partial token boundary issues, including code domains.
Magpie implications:
- Canonical
key=valueforms. - Regular opcode syntax with low branching.
- Reduced format variance across iterations.
Claim: Less ambiguous, more regular grammars are easier for model generation and correction.
Evidence:
- Constrained decoding reduces syntax errors in code generation settings.
- Grammar-constrained decoding improves syntactic correctness in structured outputs.
Magpie implications:
- Explicit opcodes, explicit call forms, no operator overloading/method-syntax ambiguity.
Claim: Format drift can alter model behavior and inflate token usage.
Evidence:
- Prompt formatting sensitivity and output variance (, ).
- Measured token savings from formatting minimization in code contexts.
- Deterministic canonicalization standards and tooling analogues (, ).
Magpie implications:
- CSNF canonical source normalization.
- Canonical JSON output for stable machine consumption.
Claim: Stable identifiers enable compact references and incremental retrieval.
Evidence:
- Content-addressed object identity in Git.
Magpie implications:
- Stable symbol-oriented references and graph artifacts.
- Better cacheability and compact cross-reference workflows.
Claim: Retrieval of relevant context is more effective than monolithic prompt stuffing.
Evidence:
- RAG , RETRO , in-context retrieval augmentation.
- LongBench’s retrieval/compression observations.
Magpie implications:
- Progressive disclosure artifacts and memory index workflows.
Magpie source files are module-centric and SSA-oriented.
- Header with strict order:
module...exports {... }imports {... }digest "..."
- Declarations (
fn,struct,enum,extern,impl,sig,global, etc.) - Function bodies are basic-block based (
bbN:labels) - Each block ends with one terminator (
ret,br,cbr,switch,unreachable)
- Explicit SSA values (
%name) - Function symbols with
@name - Type symbols with
TName - Explicit ownership forms (
shared,borrow,mutborrow,weak) - Explicit heap and projection ops (e.g.,
new,getfield,setfield)
- Line comment:
;... - Doc comment token:
;;...(double semicolon, NOT triple)
- Identifiers:
ident - Function names:
@fn_name - SSA names:
%local - Type names:
TType - Block labels:
bb0,bb1,...
- Integer literals (decimal and hex
0x...) - Float literals (
123.45, optionalf32/f64suffix) - String literals with escapes (
\n,\t,\\,\",\u{...}) - Booleans:
true,false - Unit literal form:
unit
{ } ( ) < > [ ] = : ,. ->
file := header decl*
header := "module" module_path
"exports" export_block
"imports" import_block
"digest" string_litexport_block := "{" (export_item ("," export_item)*)? "}"
export_item := fn_name | type_nameimport_block := "{" (import_group ("," import_group)*)? "}"
import_group := module_path "::" "{" (import_item ("," import_item)*)? "}"
import_item := fn_name | type_name- Canonical ordering of exports/imports.
- Canonical printing of ops and types.
- Canonical block label remapping on formatting.
- Digest normalization via
update_digest.
fn @name(...) -> Type {... }async fn @name(...) -> Type {... }unsafe fn @name(...) -> Type {... }gpu fn @name(...) -> Type target(<ident>) {... }
Optional metadata block:
meta {
uses {... }
effects {... }
cost { key=123,... }
}
heap struct TName { field... }value struct TName { field... }heap enum TName { variant... }value enum TName { variant... }
extern "c" module ffi {
fn @name(%x: i64) -> i64 attrs { link_name="...", returns="owned" }
}
global @g: i64 = const.i64 42
sig TOrdPoint(borrow TPoint, borrow TPoint) -> i32
impl ord for TPoint = @ord_point
- Signed ints:
i1 i8 i16 i32 i64 i128 - Unsigned ints:
u1 u8 u16 u32 u64 u128 - Floats:
f16 f32 f64 bf16 - Others:
bool unit
Prefix modifiers:
sharedborrowmutborrowweak
StrArray<T>Map<K, V>TOption<T>TResult<Ok, Err>TStrBuilderTMutex<T>TRwLock<T>TCell<T>TFuture<T>TChannelSend<T>TChannelRecv<T>TCallable<TSig>
- Named type:
TNameormodule.path.TName - Raw pointer:
rawptr<T>
Each basic block contains:
- SSA assignments (
%dst: Ty = value_op...) - Void ops (
setfield...,arr.push..., etc.) - Terminator (required)
ret [value]?br bbNcbr cond bbThen bbElseswitch value { case lit -> bbN... } else bbDefaultunreachable
This section lists major surface op families and canonical syntax forms.
const.<Type> <literal>
i.add,i.sub,i.mul,i.sdiv,i.udiv,i.srem,i.uremi.add.wrap,i.sub.wrap,i.mul.wrapi.add.checked,i.sub.checked,i.mul.checkedi.and,i.or,i.xor,i.shl,i.lshr,i.ashr
All use:
{ lhs=<value>, rhs=<value> }
f.add,f.sub,f.mul,f.div,f.remf.add.fast,f.sub.fast,f.mul.fast,f.div.fast
- Integer:
icmp.eq/ne/slt/sgt/sle/sge/ult/ugt/ule/uge - Float:
fcmp.oeq/one/olt/ogt/ole/oge
call @fn...call.indirect <callable>...try @fn...suspend.call @fn...suspend.await { fut=... }
new Type { field=v,... }getfield { obj=..., field=name }(keys accepted in any order)phi Type { [bb1:v1], [bb2:v2],... }
enum.new<Variant> {... }enum.tag { v=... }enum.payload<Variant> { v=... }enum.is<Variant> { v=... }
share,clone.shared,clone.weak,weak.downgrade,weak.upgradeborrow.shared,borrow.mut
cast<PrimFrom, PrimTo> { v=... }ptr.null<T>,ptr.addr<T>,ptr.from_addr<T>,ptr.add<T>,ptr.load<T>
callable.capture @fn { capture=value,... }
arr.new<T>,arr.len,arr.get,arr.pop,arr.slice,arr.contains,arr.map,arr.filter,arr.reduce
map.new<K,V>,map.len,map.get,map.get_ref,map.delete,map.contains_key,map.keys,map.values
str.concat,str.len,str.eq,str.slice,str.bytesstr.builder.new,str.builder.buildstr.parse_i64,str.parse_u64,str.parse_f64,str.parse_booljson.encode<T>,json.decode<T>
Runtime ABI note (current migration model):
- Compiler lowering uses fallible runtime calls (
mp_rt_str_try_parse_*,mp_rt_json_try_*) and checks status codes. - In the temporary compatibility path, non-OK status still routes to
mp_rt_panicto preserve legacy user-visible behavior. - Legacy runtime wrappers (
mp_rt_str_parse_*,mp_rt_json_encode/decode) are compatibility shims and are deprecated. New integrations should call only*_try_*APIs. - Ownership contract: successful
mp_rt_json_try_decodereturns an ownedout_val; release it withmp_rt_json_decoded_free(out_val, type_id).
gpu.thread_id,gpu.workgroup_id,gpu.workgroup_size,gpu.global_idgpu.buffer_load<T>,gpu.buffer_len<T>gpu.shared<count,T>gpu.launch,gpu.launch_async
call_void,call_void.indirectsetfield { obj=..., field=..., val=... }(keys accepted in any order)panic { msg=... }ptr.store<T> { p=..., v=... }arr.set,arr.push,arr.sort,arr.foreachmap.set,map.delete_voidstr.builder.append_str/i64/i32/f64/boolgpu.barriergpu.buffer_store<T>
- Single definition per local id.
- Uses must be dominated by defs.
- Branch/switch targets must exist.
- Every block has one terminator.
- Borrow values cannot be stored into escaping locations.
- Borrow handles cannot cross block boundaries.
- Borrow handles cannot appear in
phi. - Returning borrow values is forbidden.
getfieldrequiresborrow/mutborrowreceiver.setfieldrequiresmutborrowreceiver.- Collection read/write ops enforce ownership modes (read via borrow; write via unique/mutborrow).
Outside unsafe contexts, forbidden:
- raw pointer opcodes (
ptr.*) - calls to unsafe functions
arr.containsrequireseqimpl on element type.arr.sortrequiresordimpl on element type.map.new<K,V>requireshash+eqimpl forK.
- Certain aggregate/deferred forms are restricted in v0.1 checks.
suspend.callon non-function callable target forms is forbidden in v0.1.
In Magpie v0.1, TCallable exists because the language forbids closures as a primitive,
but still needs first-class “callable later” behavior (callbacks, middleware, higher-order APIs)
without hidden captures or opaque syntax.
TCallable<TSig> is an ARC-managed heap object containing:
- target function identity (call target)
- optional captured environment pointer (
data_ptr, nullable) - callable metadata/vtable entries (
call_fn,drop_fn, capture layout metadata)
Creation and invocation are explicit:
- create:
callable.capture @fn { capture1=%v1,... } - invoke:
call.indirect %callable {... }orcall_void.indirect
Magpie intentionally avoids closure syntax because closure primitives often imply:
- implicit capture set inference,
- hidden environment layout,
- hidden destructor behavior,
- increased generation ambiguity for LLM-driven tool loops.
TCallable keeps these visible and explicit.
- Visible dependencies
- Capture list is explicit in one line.
- Regular generation pattern
sig+callable.capture+call.indirectare low-ambiguity templates.
- Deterministic ownership repair
- capture is a move boundary; clone/share fixes are mechanical.
- Storage-friendly behavior objects
- Router/middleware patterns can store callables in typed fields and arrays.
- Async boundary clarity
- v0.1 forbids problematic
suspend.callcallable-indirection patterns, reducing opaque failures.
module demo.callable
exports { @main, @multiply_by }
imports { }
digest "0000000000000000"
sig TMulSig(i32) -> i32
fn @multiply_by(%x: i32, %factor: i32) -> i32 {
bb0:
%y: i32 = i.mul { lhs=%x, rhs=%factor }
ret %y
}
fn @main() -> i32 {
bb0:
%factor: i32 = const.i32 3
; Create callable with captured factor
%mul_by_3: TCallable<TMulSig> = callable.capture @multiply_by { factor=%factor }
; Invoke indirectly
%result: i32 = call.indirect %mul_by_3 { args=[const.i32 7] }
ret %result
}
Magpie intentionally combines:
- ARC for deterministic lifetime reclamation, and
- Rust-like ownership/borrowing for static alias/mutation safety.
- ARC answers: “when can this heap object be released?”
- Ownership answers: “who may alias/mutate this value right now?”
This avoids both:
- manual memory management burden, and
- unconstrained shared-mutation hazards.
- Most values stay unique
- refcount often remains near 1 unless explicitly shared.
- Explicit sharing operations
share,clone.shared, etc. make aliasing costs visible.
- Atomicity where needed
- shared/thread-crossing paths pay synchronization costs explicitly.
- Mutations require exclusive pathways (unique or mutborrow).
- Borrow restrictions prevent common temporal aliasing mistakes.
- Explicit ownership boundaries make ARC optimization passes more reliable.
%p: TPerson = new TPerson { name=%n, age=%a }
%s: shared TPerson = share { v=%p }
%s2: shared TPerson = clone.shared { v=%s }
; compiler/runtime manage release points for %s2 and %s
┌────────────┐
new -------->│ Unique │ (initial state, refcount=1)
│ (owned) │
└──┬──┬──┬───┘
│ │ │
share │ │ │ borrow.shared
(consumes) │ │ │ (temporary)
│ │ │
┌──────────────┘ │ └──────────────┐
v │ v
┌───────────┐ │ ┌───────────┐
│ Shared │ │ │ Borrow │
│ (ARC, RC) │ │ │(read-only)│
└──┬────┬───┘ │ └───────────┘
│ │ │
│ │ weak. │ borrow.mut
│ │ downgrade │ (temporary, exclusive)
│ │ │
│ ┌─v────────┐ ┌─v───────────┐
│ │ Weak │ │ MutBorrow │
│ │(non-own) │ │ (exclusive) │
│ └──────────┘ └─────────────┘
│
│ clone.shared (increments refcount)
v
┌───────────┐
│ Shared │ (additional reference)
│ (clone) │
└───────────┘
RULES:
- Borrows are block-scoped (cannot cross br/cbr boundaries)
- Borrows cannot appear in phi nodes
- Functions cannot return borrow values
- mutborrow is exclusive: no other borrows or moves while active
- share consumes the unique handle
- ARC retain/release inserted automatically by Stage 8
┌─────────────────────────────────────────────────┐
│ All Types │
├──────────────────┬──────────────────────────────┤
│ Value Types │ Heap Types │
│ (stack/inline) │ (ARC-managed handle) │
├──────────────────┼──────────────────────────────┤
│ Primitives: │ Builtins: │
│ i8..i128 │ Str │
│ u8..u128 │ Array<T> │
│ f16, f32, f64 │ Map<K, V> │
│ bf16 │ │
│ bool, unit │ TStrBuilder │
│ i1, u1 │ TFuture<T> │
│ │ TMutex<T>, TRwLock<T> │
│ Value Structs: │ TCell<T> │
│ value struct T │ TChannelSend/Recv<T> │
│ │ │
│ Value Enums: │ User Heap Types: │
│ TOption<T> │ heap struct TName │
│ TResult<O, E> │ heap enum TName │
│ │ │
│ │ Callable: │
│ │ TCallable<TSig> │
│ │ │
│ │ Pointer (unsafe): │
│ │ rawptr<T> │
└──────────────────┴──────────────────────────────┘
Ownership modifiers (apply to heap types only):
(none) = unique owned handle
shared = reference-counted (ARC)
borrow = immutable reference
mutborrow = mutable exclusive reference
weak = non-owning reference
Note: TOption and TResult are VALUE enums.
shared/weak modifiers on TOption/TResult are rejected (MPT0002/MPT0003).
Str has built-in hash, eq, ord impls (no explicit impl needed for Map keys).
High-level crate roles:
magpie_cli: command-line UX and config resolutionmagpie_driver: staged compilation orchestrationmagpie_lex: tokenizationmagpie_parse: recursive-descent parsermagpie_sema: resolve/lowering/type checks/trait checks/v0.1 checksmagpie_hir: HIR structures + verifiermagpie_own: ownership checkermagpie_mpir: MPIR + verifier + printermagpie_arc: ARC insertion/optimization passesmagpie_codegen_llvm: LLVM loweringmagpie_codegen_wasm: wasm lowering pathmagpie_mono: monomorphization (BLAKE3-keyed generic specialization)magpie_rt: runtime ABI, GPU dispatch (dlopen), profiling, MLX FFImagpie_gpu: GPU codegen core (BackendEmitter trait, CFG structurizer, kernel registry)magpie_gpu_spirv: SPIR-V backend (Vulkan)magpie_gpu_msl: Metal Shading Language backend (Apple)magpie_gpu_ptx: PTX/NVVM backend (NVIDIA CUDA)magpie_gpu_hip: HIP/HSACO backend (AMD ROCm)magpie_gpu_wgsl: WGSL backend (WebGPU)magpie_mlx: MLX host API integration (Apple Silicon ML)magpie_web: web framework + MCP integration pathsmagpie_memory: index/query workflows for memory/context artifacts
Driver stage names:
stage1_read_lex_parsestage2_resolvestage3_typecheckstage3_5_async_loweringstage4_verify_hirstage5_ownership_checkstage6_lower_mpirstage7_verify_mpirstage8_arc_insertionstage9_arc_optimizationstage10_codegenstage11_linkstage12_mms_update
This staged structure is intentionally visible in output timing and diagnostics.
.mp source
│
v
┌─────────────────────┐
│ Stage 1: Lex/Parse │──> .ast.txt
│ magpie_lex │
│ magpie_parse │
│ magpie_csnf │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 2: Resolve │ (imports, symbol tables)
│ magpie_sema │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 3: Typecheck │ (AST -> HIR, type validation)
│ magpie_sema │
│ magpie_types │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 3.5: Async │ (coroutine state machines)
│ Lowering │ suspend.call -> dispatch switch
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 4: Verify HIR │ (SSA, borrow invariants)
│ magpie_hir │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 5: Ownership │ (move/borrow/alias rules)
│ magpie_own │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 6: Lower MPIR │──> .mpir
│ magpie_mpir │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 7: Verify MPIR│ (SID/CFG/type invariants)
│ magpie_mpir │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 8: ARC Insert │ (retain/release insertion)
│ magpie_arc │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 9: ARC Opt │ (elide redundant refcounting)
│ magpie_arc │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 10: Codegen │──> .ll (LLVM IR), .gpu_registry.ll
│ magpie_codegen_llvm│ GPU backends: .spv, .metal, .ptx, .hip, .wgsl
│ magpie_gpu + 5 │ (spirv/msl/ptx/hip/wgsl emitters)
│ magpie_mono │ Monomorphization: generic specialization
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 11: Link │──> native executable / shared lib
│ clang -x ir │ (or lli for interpretation)
│ libmagpie_rt.a │
└──────────┬──────────┘
v
┌─────────────────────┐
│ Stage 12: MMS Update│──> .mms_index.json
│ magpie_memory │
└─────────────────────┘
magpie_cli
│
magpie_driver ─────────────────────────────┐
/ │ \ │
magpie_lex magpie_sema magpie_codegen_llvm magpie_web
│ │ │ │ │
magpie_parse magpie_hir magpie_arc magpie_jit
│ │ │
magpie_ast magpie_own magpie_mpir ──── magpie_mono
│ │ │
magpie_csnf magpie_types magpie_gpu
│ ├── magpie_gpu_spirv (Vulkan)
magpie_diag ├── magpie_gpu_msl (Metal)
├── magpie_gpu_ptx (CUDA)
├── magpie_gpu_hip (ROCm)
└── magpie_gpu_wgsl (WebGPU)
magpie_rt (runtime, GPU dispatch, profiling)
magpie_mlx (MLX Apple Silicon ML)
magpie_pkg
magpie_memory
magpie_ctx
Supported emit kinds include:
llvm-irllvm-bcobjectasmexeshared-libmpirmpdmpdbgsymgraphdepsgraphownershipgraphcfggraph
GPU backend emit kinds:
spv— SPIR-V binary module (Vulkan)msl— Metal Shading Language text (Apple)ptx— PTX via LLVM IR with nvptx64 triple (NVIDIA CUDA)hip— HSACO via LLVM IR with amdgcn triple (AMD ROCm)wgsl— WGSL text (WebGPU)
Typical usage:
magpie build --entry src/main.mp --emit mpir,llvm-ir,mpdbgMagpie diagnostics are structured objects with:
- code (
MPS0001,MPT2014, etc.) - severity
- message/title
- spans
- optional explanation / fix hints
- optional trace/rag/doc links
Common code families:
MPP*parse/io/artifactMPS*resolve/SSA/structural invariantsMPT*type/trait/v0.1 restrictionsMPO*ownership/borrowingMPF*FFIMPG*GPU (39 codes:MPG_TYP_*type,MPG_KRN_*kernel,MPG_BUF_*buffer,MPG_SYN_*sync,MPG_CAP_*capability,MPG_LNK_*link,MPG_MLX_*MLX,MPG_PRF_*profiling)MPL*lint/link/LLM budgetMPW*webMPK*package/dependencyMPM*memory/index
Parse/JSON sema diagnostics (migration-focused):
| Code | Meaning | Trigger |
|---|---|---|
MPT2033 |
Parse/JSON result shape mismatch | Result type is neither legacy shape nor TResult<ok, err> shape expected by the opcode |
MPT2034 |
Parse/JSON input type mismatch | Parse/decode input is unknown or not Str / borrow Str |
MPT2035 |
json.encode<T> value type mismatch |
Encoded value type does not match generic target T (or value type is unknown) |
Explain command:
magpie explain MPT2014 --output jsonThis section is an implementation-level argument reference.
| Flag | Type | Default | Notes |
|---|---|---|---|
--output |
enum | text |
text, json, jsonl |
--color |
enum | auto |
auto, always, never |
--log-level |
enum | warn |
error, warn, info, debug, trace |
--profile |
enum | dev |
CLI parser accepts dev, release, custom; config maps non-release to dev |
--target |
string | host default | target triple |
--emit |
csv string | command-dependent | artifact kinds |
--entry |
path | manifest/default | source entry file |
--cache-dir |
path | none | cache path |
-j, --jobs |
int | none | parallel jobs |
--features |
csv string | empty | feature flags |
--no-default-features |
bool | false | disable default features |
--offline |
bool | false | dependency operations offline |
--llm |
bool | false | LLM-optimized output mode |
--no-auto-fmt |
bool | false | disable pre-build auto-format in llm mode |
--llm-token-budget |
int | resolved | output budget |
--llm-tokenizer |
string | resolved | tokenizer id |
--llm-budget-policy |
enum | resolved | balanced, diagnostics_first, slices_first, minimal |
--max-errors |
int | 20 |
max diagnostics per pass |
--shared-generics |
bool | false | use shared generics mode |
magpie new <name>magpie buildmagpie run [args...]magpie replmagpie fmt [--fix-meta]magpie parse [--emit ast]magpie lintmagpie test [--filter <pattern>]magpie docmagpie mpir verifymagpie explain <CODE>magpie pkg resolve|add|remove|whymagpie web dev|build|servemagpie mcp servemagpie memory build|query --q <query> [--k 10]magpie ctx packmagpie ffi import --header <h> --out <mp>magpie graph symbols|deps|ownership|cfg
--fix-meta
--emit ast(currently fixed toast)
--filter <pattern>
-q, --q <query>-k, --k <top_k>(default 10)
--header <header-path>--out <output-path>
rundefault emit:- release profile:
exe - dev profile:
llvm-ir testmode auto-addstestfeature when absent.- In
--llmmode (unless--no-auto-fmt), auto-format precheck runs first.
Selected settings are resolved from a combination of:
- CLI flags
- Environment variables (
MAGPIE_LLM,MAGPIE_LLM_TOKEN_BUDGET) Magpie.tomldefaults ([build],[llm])- hardcoded defaults
[build].entry[llm].mode_default[llm].token_budget[llm].tokenizer[llm].budget_policy[gpu].backend—auto,spirv,msl,ptx,hip,wgsl[gpu].fallback—cpu(default) orerror[gpu].llc_path— custom path tollcbinary (PTX/HIP backends)[gpu].lld_path— custom path told.lldbinary (HIP backend)
magpie build --entry src/main.mp --emit mpir,llvm-ir --output jsonmagpie run --profile release --entry src/main.mp --emit exemagpie parse --entry src/main.mp --output jsonmagpie graph symbols --entry src/main.mp --output json
magpie graph deps --entry src/main.mp --output json
magpie graph ownership --entry src/main.mp --output json
magpie graph cfg --entry src/main.mp --output jsonmagpie build --entry src/main.mp --output json --emit mpir,llvm-ir,mpdbg
magpie explain MPS0024 --output jsonmodule demo.main
exports { @main }
imports { }
digest "0000000000000000"
fn @main() -> i32 {
bb0:
ret const.i32 0
}
module demo.point
exports { @main }
imports { }
digest "0000000000000000"
heap struct TPoint {
field x: i64
field y: i64
}
fn @main() -> i64 {
bb0:
%p: TPoint = new TPoint { x=const.i64 1, y=const.i64 2 }
%pm: mutborrow TPoint = borrow.mut { v=%p }
setfield { obj=%pm, field=y, val=const.i64 3 }
br bb1
bb1:
%pb: borrow TPoint = borrow.shared { v=%p }
%y: i64 = getfield { obj=%pb, field=y }
ret %y
}
module demo.callable
exports { @main, @multiply_by }
imports { }
digest "0000000000000000"
sig TMulSig(i32) -> i32
fn @multiply_by(%x: i32, %factor: i32) -> i32 {
bb0:
%y: i32 = i.mul { lhs=%x, rhs=%factor }
ret %y
}
fn @main() -> i32 {
bb0:
%factor: i32 = const.i32 3
%mul_by_3: TCallable<TMulSig> = callable.capture @multiply_by { factor=%factor }
%result: i32 = call.indirect %mul_by_3 { args=[const.i32 7] }
ret %result
}
- This document is intentionally comprehensive and practical.
- For binary-only operation, prioritize
--output json,magpie explain <CODE>, and emitted artifacts. - For source-level implementation changes, use stage-specific diagnostics and crate boundaries to localize fixes.
The following is an extended grammar sketch aligned with v0.1 parser behavior.
file := header decl*
header := module_decl exports_decl imports_decl digest_decl
module_decl := "module" module_path
exports_decl := "exports" "{" export_item_list? "}"
imports_decl := "imports" "{" import_group_list? "}"
digest_decl := "digest" string_lit
export_item_list := export_item ("," export_item)*
export_item := fn_name | type_name
import_group_list := import_group ("," import_group)*
import_group := module_path "::" "{" import_item_list? "}"
import_item_list := import_item ("," import_item)*
import_item := fn_name | type_name
decl := fn_decl
| async_fn_decl
| unsafe_fn_decl
| gpu_fn_decl
| heap_struct_decl
| value_struct_decl
| heap_enum_decl
| value_enum_decl
| extern_decl
| global_decl
| impl_decl
| sig_decl
fn_decl := doc* "fn" fn_name "(" params? ")" "->" type fn_meta? blocks
async_fn_decl := doc* "async" "fn" fn_name "(" params? ")" "->" type fn_meta? blocks
unsafe_fn_decl := doc* "unsafe" "fn" fn_name "(" params? ")" "->" type fn_meta? blocks
gpu_fn_decl := doc* "gpu" "fn" fn_name "(" params? ")" "->" type "target" "(" ident ")" fn_meta? blocks
fn_meta := "meta" "{" (meta_uses | meta_effects | meta_cost)* "}"
meta_uses := "uses" "{" fqn_list? "}"
meta_effects := "effects" "{" ident_list? "}"
meta_cost := "cost" "{" kv_i64_list? "}"
params := param ("," param)*
param := ssa_name ":" type
heap_struct_decl := doc* "heap" "struct" type_name type_params? "{" field_decl* "}"
value_struct_decl := doc* "value" "struct" type_name type_params? "{" field_decl* "}"
field_decl := "field" ident ":" type
heap_enum_decl := doc* "heap" "enum" type_name type_params? "{" variant_decl* "}"
value_enum_decl := doc* "value" "enum" type_name type_params? "{" variant_decl* "}"
variant_decl := "variant" ident "{" field_decl_inline_list? "}"
extern_decl := doc* "extern" string_lit "module" ident "{" extern_item* "}"
extern_item := "fn" fn_name "(" params? ")" "->" type attrs_block?
attrs_block := "attrs" "{" kv_string_list? "}"
global_decl := doc* "global" fn_name ":" type "=" const_expr
impl_decl := "impl" ident "for" type "=" fn_ref
sig_decl := "sig" type_name "(" type_list? ")" "->" type
blocks := "{" block+ "}"
block := block_label ":" instr* terminator
instr := assign_instr
| void_instr
| unsafe_block
assign_instr := ssa_name ":" type "=" value_op
void_instr := void_op
unsafe_block := "unsafe" "{" (assign_instr | void_instr)+ "}"
terminator := "ret" value_ref?
| "br" block_label
| "cbr" value_ref block_label block_label
| "switch" value_ref "{" switch_arms* "}" "else" block_label
| "unreachable"
switch_arms := "case" const_lit "->" block_label
value_ref := ssa_name | const_expr
const_expr := "const" "." type const_lit
type := ownership_mod? base_type
ownership_mod := "shared" | "borrow" | "mutborrow" | "weak"
base_type := prim_type
| builtin_type
| named_type
| rawptr_type
| callable_type
rawptr_type := "rawptr" "<" type ">"
callable_type := "TCallable" "<" type_ref ">"
named_type := (module_path ".")? type_name type_args?
type_args := "<" type_list ">"
type_params := "<" type_param_list ">"
type_param_list := type_param ("," type_param)*
type_param := ident ":" ident
module_path := ident ("." ident)*
fn_ref := fn_name | module_path "." fn_name
type_ref := type_name | module_path "." type_name
prim_type := "i1" | "i8" | "i16" | "i32" | "i64" | "i128"
| "u1" | "u8" | "u16" | "u32" | "u64" | "u128"
| "f16" | "f32" | "f64" | "bf16"
| "bool" | "unit"
builtin_type := "Str"
| "Array" "<" type ">"
| "Map" "<" type "," type ">"
| "TOption" "<" type ">"
| "TResult" "<" type "," type ">"
| "TStrBuilder"
| "TMutex" "<" type ">"
| "TRwLock" "<" type ">"
| "TCell" "<" type ">"
| "TFuture" "<" type ">"
| "TChannelSend" "<" type ">"
| "TChannelRecv" "<" type ">"i.add { lhs=V, rhs=V }
i.sub { lhs=V, rhs=V }
i.mul { lhs=V, rhs=V }
i.sdiv { lhs=V, rhs=V }
i.udiv { lhs=V, rhs=V }
i.srem { lhs=V, rhs=V }
i.urem { lhs=V, rhs=V }
i.add.wrap { lhs=V, rhs=V }
i.sub.wrap { lhs=V, rhs=V }
i.mul.wrap { lhs=V, rhs=V }
i.add.checked { lhs=V, rhs=V }
i.sub.checked { lhs=V, rhs=V }
i.mul.checked { lhs=V, rhs=V }
i.and { lhs=V, rhs=V }
i.or { lhs=V, rhs=V }
i.xor { lhs=V, rhs=V }
i.shl { lhs=V, rhs=V }
i.lshr { lhs=V, rhs=V }
i.ashr { lhs=V, rhs=V }
f.add { lhs=V, rhs=V }
f.sub { lhs=V, rhs=V }
f.mul { lhs=V, rhs=V }
f.div { lhs=V, rhs=V }
f.rem { lhs=V, rhs=V }
f.add.fast { lhs=V, rhs=V }
f.sub.fast { lhs=V, rhs=V }
f.mul.fast { lhs=V, rhs=V }
f.div.fast { lhs=V, rhs=V }
icmp.eq { lhs=V, rhs=V }
icmp.ne { lhs=V, rhs=V }
icmp.slt { lhs=V, rhs=V }
icmp.sgt { lhs=V, rhs=V }
icmp.sle { lhs=V, rhs=V }
icmp.sge { lhs=V, rhs=V }
icmp.ult { lhs=V, rhs=V }
icmp.ugt { lhs=V, rhs=V }
icmp.ule { lhs=V, rhs=V }
icmp.uge { lhs=V, rhs=V }
fcmp.oeq { lhs=V, rhs=V }
fcmp.one { lhs=V, rhs=V }
fcmp.olt { lhs=V, rhs=V }
fcmp.ogt { lhs=V, rhs=V }
fcmp.ole { lhs=V, rhs=V }
fcmp.oge { lhs=V, rhs=V }
call @fn<TypeArgs?> { key=Arg,... }
call.indirect V { key=Arg,... }
try @fn<TypeArgs?> { key=Arg,... }
suspend.call @fn<TypeArgs?> { key=Arg,... }
suspend.await { fut=V }
new Type { field=V,... }
getfield { obj=V, field=name }
phi Type { [bbN:V], [bbM:V],... }
enum.new<Variant> { key=V,... }
enum.tag { v=V }
enum.payload<Variant> { v=V }
enum.is<Variant> { v=V }
share { v=V }
clone.shared { v=V }
clone.weak { v=V }
weak.downgrade { v=V }
weak.upgrade { v=V }
cast<PrimFrom, PrimTo> { v=V }
borrow.shared { v=V }
borrow.mut { v=V }
ptr.null<T>
ptr.addr<T> { p=V }
ptr.from_addr<T> { addr=V }
ptr.add<T> { p=V, count=V }
ptr.load<T> { p=V }
callable.capture @fn { cap_name=V,... }
arr.new<T> { cap=V }
arr.len { arr=V }
arr.get { arr=V, idx=V }
arr.pop { arr=V }
arr.slice { arr=V, start=V, end=V }
arr.contains { arr=V, val=V }
arr.map { arr=V, fn=V }
arr.filter { arr=V, fn=V }
arr.reduce { arr=V, init=V, fn=V }
map.new<K, V> { }
map.len { map=V }
map.get { map=V, key=V }
map.get_ref { map=V, key=V }
map.delete { map=V, key=V }
map.contains_key { map=V, key=V }
map.keys { map=V }
map.values { map=V }
str.concat { a=V, b=V }
str.len { s=V }
str.eq { a=V, b=V }
str.slice { s=V, start=V, end=V }
str.bytes { s=V }
str.builder.new { }
str.builder.build { b=V }
str.parse_i64 { s=V }
str.parse_u64 { s=V }
str.parse_f64 { s=V }
str.parse_bool { s=V }
json.encode<T> { v=V }
json.decode<T> { s=V }
gpu.thread_id { dim=V }
gpu.workgroup_id { dim=V }
gpu.workgroup_size { dim=V }
gpu.global_id { dim=V }
gpu.buffer_load<T> { buf=V, idx=V }
gpu.buffer_len<T> { buf=V }
gpu.shared<count, T>
gpu.launch { device=V, kernel=@fn, grid=Arg, block=Arg, args=Arg }
gpu.launch_async { device=V, kernel=@fn, grid=Arg, block=Arg, args=Arg }
Compatibility note:
- The source op names above are stable.
- Internally, parse/json codegen now targets
*_try_*runtime symbols with explicit status branching at the ABI boundary.
call_void @fn<TypeArgs?> { key=Arg,... }
call_void.indirect V { key=Arg,... }
setfield { obj=V, field=name, val=V }
panic { msg=V }
ptr.store<T> { p=V, v=V }
arr.set { arr=V, idx=V, val=V }
arr.push { arr=V, val=V }
arr.sort { arr=V }
arr.foreach { arr=V, fn=V }
map.set { map=V, key=V, val=V }
map.delete_void { map=V, key=V }
str.builder.append_str { b=V, s=V }
str.builder.append_i64 { b=V, v=V }
str.builder.append_i32 { b=V, v=V }
str.builder.append_f64 { b=V, v=V }
str.builder.append_bool { b=V, v=V }
gpu.barrier
gpu.buffer_store<T> { buf=V, idx=V, v=V }
magpie [GLOBAL_FLAGS] <command> [SUBCOMMAND_FLAGS]magpie --output json --entry src/main.mp build
magpie --profile release --target x86_64-unknown-linux-gnu --emit exe run
magpie --llm --llm-token-budget 12000 --llm-budget-policy balanced build
magpie --features test,web --no-default-features test --filter callablemagpie new demo_project
magpie fmt --fix-meta
magpie parse --entry src/main.mp --emit ast
magpie lint --entry src/main.mp
magpie doc
magpie explain MPO0102
magpie mpir verify --entry src/main.mp
magpie graph symbols --entry src/main.mp
magpie graph deps --entry src/main.mp
magpie graph ownership --entry src/main.mp
magpie graph cfg --entry src/main.mp
magpie ffi import --header ffi.h --out ffi_bindings.mp
magpie pkg resolve
magpie pkg add serde_like
magpie pkg remove serde_like
magpie pkg why std
magpie web dev
magpie web build
magpie web serve
magpie mcp serve
magpie memory build --entry src/main.mp
magpie memory query -q "ownershipgraph borrow phi" -k 10 --entry src/main.mp
magpie ctx pack --entry src/main.mp| Claim | Operational basis in Magpie | Confidence |
|---|---|---|
| Bounded output design is necessary | Token-budget options, JSON envelopes, and progressive emit strategy | High |
| Locality improves practical reliability | Explicit control flow, explicit ownership ops, explicit call forms | High |
| Canonical formatting improves stability | CSNF formatting and deterministic output behavior | High |
| Progressive disclosure outperforms bulk dumps in workflows | mpdbg, graph emits, and memory/query workflows |
High |
| Stable symbol identity helps incremental workflows | Symbol/dependency graphs and deterministic artifact surfaces | Medium-High |
| Canonicalization reduces iterative edit drift | Deterministic formatting and reduced syntactic variance | Medium |
Magpie supports 5 GPU compute backends via a unified BackendEmitter trait defined in magpie_gpu:
pub trait BackendEmitter {
fn backend_id(&self) -> GpuBackend;
fn validate_kernel(&self, func: &MpirFn) -> Result<(), String>;
fn compute_layout(&self, func: &MpirFn, type_ctx: &TypeCtx) -> KernelLayout;
fn emit_kernel(&self, func: &MpirFn, type_ctx: &TypeCtx) -> Result<Vec<u8>, String>;
fn artifact_extension(&self) -> &'static str;
}| Backend | Enum | Target triple | Control flow | Output |
|---|---|---|---|---|
| SPIR-V | GpuBackend::Spv (1) |
— | Unstructured CFG | Binary SPIR-V module |
| MSL | GpuBackend::Msl (2) |
— | Structurized (Relooper) | Metal Shading Language text |
| PTX | GpuBackend::Ptx (3) |
nvptx64-nvidia-cuda |
Unstructured LLVM IR | PTX via llc |
| HIP | GpuBackend::Hip (4) |
amdgcn-amd-amdhsa |
Unstructured LLVM IR | HSACO via llc + ld.lld |
| WGSL | GpuBackend::Wgsl (5) |
— | Structurized (Relooper) | WGSL text |
MSL and WGSL require structured control flow (no arbitrary goto). The magpie_gpu::structurize module implements a Relooper-style algorithm that converts MPIR basic blocks into structured nodes:
StructuredNode::Block— sequential codeStructuredNode::IfElse— conditional branchStructuredNode::Loop— loop with break/continueStructuredNode::Return— function return
The runtime (magpie_rt) probes GPU backends at startup via dlopen in priority order:
- Metal.framework (macOS)
- libcuda.so / nvcuda.dll (NVIDIA)
- libamdhip64.so (AMD ROCm)
- libvulkan.so / vulkan-1.dll (Vulkan)
- libwgpu_native.so (WebGPU)
Falls back to CPU simulation if no GPU backend is available.
9 runtime ABI functions for GPU performance analysis:
- Session lifecycle:
mp_rt_gpu_profile_begin()/mp_rt_gpu_profile_end() - Region markers:
mp_rt_gpu_profile_mark_begin()/mp_rt_gpu_profile_mark_end() - Export:
mp_rt_gpu_profile_export_chrome()— Chrome trace JSON (chrome://tracing) - Hardware counters:
mp_rt_gpu_profile_available_counters()/enable_counters()/read_counters() - Memory:
mp_rt_gpu_profile_memory_stats()— allocation tracking (bytes/peak/count)
[gpu]
backend = "auto" # auto | spirv | msl | ptx | hip | wgsl
fallback = "cpu" # cpu | error
llc_path = "/path/to/llc" # optional: custom llc for PTX/HIP
lld_path = "/path/to/ld.lld" # optional: custom lld for HIP HSACO| TypeId | Type |
|---|---|
| 16 | bf16 (bfloat16 primitive) |
| 33 | gpu.TError |
| 34 | gpu.TErrorKind |
| 35 | gpu.TProfileSession |
| 36 | gpu.TProfileEvent |
| 37 | gpu.TMemoryStats |
| 38 | mlx.TArrayBase |
| 39 | mlx.TLayerHandle |
| 40 | mlx.TOptimizerHandle |
| 50 | gpu.TKernelInternal |
39 diagnostic codes in the MPG_XXX_NNNN format covering:
MPG_TYP_*— type errors (bf16 unsupported on backend, invalid buffer element type)MPG_KRN_*— kernel validation (workgroup size limits, buffer count limits)MPG_BUF_*— buffer errors (type mismatch, out of bounds)MPG_SYN_*— synchronization (barrier misuse, fence errors)MPG_CAP_*— capability (missing backend feature, unsupported op)MPG_LNK_*— link errors (llc/lld not found, compilation failure)MPG_MLX_*— MLX dispatch (dlopen failure, missing symbol)MPG_PRF_*— profiling (session not started, export failure)
The magpie_mlx crate provides full MLX host API integration via dlopen/dlsym dispatch tables (~40 function pointers). MLX is loaded at runtime, not at build time, so the compiler can be built on any platform.
Required symbols (fail if missing): array lifecycle, element-wise ops, shape, reduction, eval, nn, random, optim, grad.
Optional symbols (graceful degradation): transpose, expand_dims, squeeze, neg, abs, exp, log, sqrt, pow, max, min, argmax, argmin, linalg (norm, inv), nn activations (relu, gelu, softmax, layer_norm, batch_norm), optimizers (sgd, adamw), value_and_grad.
50+ mp_rt_mlx_* functions exposed for codegen integration:
- Array:
array_from_data,array_shape,array_eval,array_to_data - Ops:
add,subtract,multiply,divide,matmul - NN:
linear,conv2d,relu,gelu,softmax,layer_norm,batch_norm - Optim:
sgd_step,adamw_step - Autograd:
grad,value_and_grad
The magpie_mono crate implements generic function specialization.
- Walk the MPIR module to find generic function calls
- For each unique concrete type argument set, compute a BLAKE3 hash
- Generate a specialized SID:
original_sid$mono$<hash> - Duplicate the function body with type parameters replaced by concrete types
- Rewrite call sites to target the specialized SID
monomorphize(modules, type_ctx)— main entry point, processes all modulesspecialize_function(func, type_args, type_ctx)— specialize a single functionsubstitute_type(ty, bindings)— replace type parameters with concrete types