Last updated: 2026-09-26
Schemas and Scenarios: A Working Z/BDD Hybrid
Formal Methods covers a Z schema: a universally-quantified invariant, checkable independently of any one concrete case. BDD as Specification covers the opposite kind of claim: a Given/When/Then scenario is one concrete, witnessed example, useful precisely because it’s concrete, not despite it. This page connects them for real: a Z-style schema states a general invariant, and a set of Gherkin-style scenarios serve as its concrete witnesses, checked against each other rather than written and trusted independently. Every mechanism described below is built, tested, and runnable — the live examples further down run in this page, not on a build machine somewhere else. The libraries themselves are in the project’s GitHub repository, named zs_*.patlang.
Prior Art
This isn’t unclaimed territory. Bowen Liu’s research at the University of Waikato converts BDD-style behavioural specifications into first-order-logic predicates and checks them for consistency against formal models1. The follow-on PhD, supervised by Judy Bowen, Jessica Turner, and Steve Reeves, does this against Z specifications and applies it to the T34 syringe driver, a safety-critical medical device2. That thesis experiments with refinement of the Z specification and leaves it as open future work2.
Liu’s process is manual and works across two documents. An analyst extracts the essential facts from a feature’s scenarios, finds the matching statements in the Z specification, writes the Given and When conditions as first-order predicates, and states each as a temporal-logic assertion: in every state, if the predicates do not all hold, the operation is not enabled. The ProB model checker tests the assertion over the Z model and returns a counter-example when it fails2. The thesis states that the process finds inconsistencies between the two specifications and does not guarantee the correctness of the system2.
The idea of checking a Gherkin specification against a Z-style specification is Liu’s. Treating the scenarios as concrete witnesses of a schema, with a tagged diagnosis for each way a witness can fail, is our own formulation, and its mechanism is different. A schema here is executable PatLang held in the same source as its scenarios. One run checks each scenario’s before and after states against the invariant, the operation’s precondition, and its postcondition, and reports which of those four checks failed. Liu’s predicates come from the Given and When clauses, and the thesis excludes negative scenarios2, so postconditions and rejection scenarios sit outside that process. Running a synthesized function or a planner’s output against a schema has no counterpart in the thesis either.
A third piece of work sits between the other two. Lili Shao’s dissertation, also from Waikato and also supervised by Judy Bowen — who shared a copy of it directly — converts a real Z specification into Gherkin and then into Playwright scripts using an LLM at each step, run against a real deployed web system4. It is not yet listed in the university’s research repository. Where Liu formalises the Gherkin and checks it against Z with ProB, Shao trusts the LLM’s translation once it parses and reads plausibly, verifying it by a Gherkin-syntax check plus manually counting how much of the Z specification’s scenarios and assertions the generated code covers. Nothing checks that the generated Gherkin means what the Z specification says. Her second experiment names the same problem as this page’s live examples: neither Z nor Gherkin can predict the name of a real page element, so the generated Playwright tests failed entirely until the selectors were fixed by hand. She calls this the mapping problem and leaves a declarative mapping language between abstract concepts and concrete implementation elements as future work. Her own tests do something none of the mechanisms on this page do: they run against a deployed system and catch defects a static check cannot. Ours checks whether a scenario is true of the schema; hers checks whether the generated code executes.
The four live examples near the top of this page check the states their scenarios supply. A second layer, described under From Examples to Every Reachable State, declares a schema as text and checks it against every state it can reach, reproduces the check from Liu’s thesis on the thesis’s own vending machine, and runs live in the panel further down. Three differences remain in Liu’s favour:
- Liu’s process is written against Z. The declared schemas use an ASCII notation with sets, maps, and quantifiers, and have no schema calculus and no unbounded types.
- Liu’s analyst finds the statements in the Z specification that match each phrase in a scenario. Here the analyst writes that correspondence down as fact and goal lines, and the check runs from there.
- Exploration covers the states reachable from one initial state through finite input domains that the schema declares, and it needs each operation’s
ensureclauses to define its after-state. ProB works over the Z model itself.
Two Kinds of Claim, Briefly FoundationalKnowledge that endures for decades — core principles
A Z schema states what must hold for every input satisfying its precondition — a BorrowBook schema states that any title already in books and not already in borrowedBy leads to a specific update, for every such title and member, not just one. A BDD scenario states one specific, concrete case; BDD as Specification is explicit that this is deliberate — a scenario is training data and a pass/fail oracle precisely because it’s concrete. Neither substitutes for the other. That’s exactly why keeping them consistent with each other is worth doing rather than assuming it happens for free.
Design notes on two choices behind this mechanism — why a schema’s clauses are ordinary compiled functions rather than inline text, and a checklist of what each stage needed — are on a separate page.
Try It Yourself: The Mechanism, Live
The four examples below run entirely in this page — the same self-hosted PatLang compiler and interpreter compiled to WebAssembly that powers every other live demo on this site, with the new pset/pmap/schema_bdd libraries inlined directly into the source (there’s no real filesystem in the browser sandbox, so include isn’t available here — everything the example needs is in the one block of text below). The last two examples carry more with them than the first two: the planner example needs plan_with_state, and the inductive-synthesis example includes the lexer, parser and lowerer as well, because the induction engine compiles each candidate rule into a small test script and runs it. Pick an example from the dropdown, edit it if you like, and press Run.
# =============================================================================
# Test framework (Stage 1 dialect): unit assertions plus a Gherkin-style
# feature runner. Step definitions are registered in the object store keyed
# by their text; features are plain text dispatched line by line, so the
# same framework covers unit, integration, and behaviour tests.
# =============================================================================
make a function called t_init returns done
set_var("t_pass", 0)
set_var("t_fail", 0)
set_var("t_tagfilter", "")
set_var("t_pending_tags", "")
set_var("t_skipping", 0)
return true
end
make a function called contains_text takes hay, needle returns r
if needle.length > hay.length then
return false
end
let i = 0
while i <= hay.length - needle.length do
if substr(hay, i, needle.length) == needle then
return true
end
let i = i + 1
end
return false
end
make a function called check takes label, actual, expected returns ok
if actual == expected then
set_var("t_pass", get("__vars", "t_pass") + 1)
print(" ok: " + label)
return true
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: " + label + " (got " + actual + ", want " + expected + ")")
return false
end
end
make a function called t_report returns ok
let p = get("__vars", "t_pass")
let f = get("__vars", "t_fail")
print("tests: " + p + " passed, " + f + " failed")
if f == 0 then
print("ALL TESTS PASSED")
return true
else
print("TESTS FAILED")
return false
end
end
# ---- Gherkin runner ----
# Register a step: step("a fresh till", "st_fresh_till")
make a function called step takes text, fname returns done
new("Step", text)
send(text, "set", "fn", fname)
return true
end
make a function called starts_with takes s, prefix returns r
if s.length < prefix.length then
return false
end
return substr(s, 0, prefix.length) == prefix
end
make a function called trim_left takes s returns out
let i = 0
let scanning = true
while (i < s.length) and scanning do
let c = char_code(s, i)
if (c == 32) or (c == 9) then
let i = i + 1
else
let scanning = false
end
end
return substr(s, i, s.length - i)
end
# Strip a Gherkin keyword; returns the step text or "" if not a step line
make a function called step_text takes line returns out
if starts_with(line, "Given ") then
return substr(line, 6, line.length - 6)
end
if starts_with(line, "When ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "Then ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "And ") then
return substr(line, 4, line.length - 4)
end
return ""
end
# Run only scenarios whose preceding @tag line contains `tag` ("" = all)
make a function called run_feature_tagged takes feature, tag returns ok
set_var("t_tagfilter", tag)
return run_feature(feature)
end
make a function called run_feature_file takes path returns ok
return run_feature(read_file(path))
end
make a function called run_feature takes feature returns ok
let h = str_intern(feature)
let n = sc_len(h)
let i = 0
let line = sb_new()
while i <= n do
let c = sc_code(h, i)
if (c == 10) or (c == -1) then
let raw = trim_left(sb_str(line))
let line = sb_new()
if starts_with(raw, "@") then
set_var("t_pending_tags", raw)
end
if starts_with(raw, "Feature:") then
print(raw)
end
if starts_with(raw, "Scenario:") then
let filter = get("__vars", "t_tagfilter")
let tags = get("__vars", "t_pending_tags")
set_var("t_pending_tags", "")
if filter then
if tags then
if contains_text(tags, filter) then
set_var("t_skipping", 0)
print(raw + " [" + tags + "]")
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 0)
print(raw)
end
else
let text = step_text(raw)
if (text != "") and (get("__vars", "t_skipping") != 1) then
if starts_with(text, "require ") then
handle_contract_step("require", substr(text, 8, text.length - 8))
else
if starts_with(text, "ensure ") then
handle_contract_step("ensure", substr(text, 7, text.length - 7))
else
let fname = get(text, "fn")
if fname then
apply(fname)
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: undefined step: " + text)
end
end
end
end
end
let i = i + 1
else
if c == 13 then
let i = i + 1
else
sb_push(line, sc_char(h, i))
let i = i + 1
end
end
end
return true
end
# =============================================================================
# pset: a minimal Set<T> built on a plain list with linear-scan membership,
# the same idiom already used by depgraph_list_contains and
# synth5_list_contains (self_hosting/lib/depgraph.patlang,
# self_hosting/lib/synthesis_lgg.patlang) rather than a new primitive --
# PatLang has no native Set type.
#
# Functional / return-new-collection style throughout (mirrors list_push's
# own semantics): every mutating-sounding operation returns a NEW list
# rather than mutating in place. This matters specifically for schema_bdd's
# before/after snapshotting, which needs an independent copy of state at
# two points in time -- an in-place structure would make "before" silently
# turn into "after" once the operation runs.
# =============================================================================
make a function called pset_new returns s
return []
end
make a function called pset_contains takes s, item returns found
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] == item then
return true
end
let i = i + 1
end
return false
end
make a function called pset_add takes s, item returns s2
if pset_contains(s, item) then
return s
end
return list_push(s, item)
end
make a function called pset_remove takes s, item returns s2
let out = []
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] != item then
let out = list_push(out, s[i])
end
let i = i + 1
end
return out
end
make a function called pset_size takes s returns n
return to_num(list_len(s))
end
# Identity -- exposed so callers iterating a set's members don't need to
# know it's "just a list" internally; the name documents intent at the
# call site instead.
make a function called pset_to_list takes s returns items
return s
end
# Set equality: same size, and every member of a is in b. Given pset_add's
# own dedup, equal size plus one-directional containment implies the
# reverse containment too (no way for b to hold something a doesn't
# without differing in size).
make a function called pset_equal takes a, b returns eq
if pset_size(a) != pset_size(b) then
return false
end
let i = 0
let n = to_num(list_len(a))
while i < n do
if pset_contains(b, a[i]) == false then
return false
end
let i = i + 1
end
return true
end
# =============================================================================
# pmap: a minimal Map<K,V> as an association list of [key, value] pairs with
# linear-scan lookup -- the exact same idiom already proven in
# self_hosting/lib/depgraph.patlang's depgraph_map_get/depgraph_map_append
# (a file->list-of-files map), generalised here to an arbitrary value type
# and given a full put/get/has/remove/keys surface. PatLang has no native
# Map/Dict type usable from self-hosted code at this scale.
#
# Functional / return-new-collection style throughout, for the same reason
# as pset.patlang: schema_bdd needs an independent before/after snapshot of
# state, which an in-place-mutating map wouldn't give for free.
#
# pmap_get returns [] (an empty list) for a missing key -- the same
# not-found sentinel depgraph_map_get already uses, not a distinct "Unit"
# literal (PatLang's dialect has no source-level Unit/nil literal to write
# directly). This is INHERENTLY AMBIGUOUS if a real stored value could
# itself be an empty list: pmap_has is the only call that actually
# distinguishes "absent" from "present but happens to look like the
# sentinel" -- never infer presence from pmap_get's return value alone.
# =============================================================================
make a function called pmap_new returns m
return []
end
make a function called pmap_has takes m, key returns found
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return true
end
let i = i + 1
end
return false
end
# See the module-level warning above: [] means "not found" here, which is
# indistinguishable from a genuinely stored empty-list value. Call
# pmap_has first whenever that distinction matters.
make a function called pmap_get takes m, key returns value
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return m[i][1]
end
let i = i + 1
end
return []
end
make a function called pmap_put takes m, key, value returns m2
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return list_set(m, i, [key, value])
end
let i = i + 1
end
return list_push(m, [key, value])
end
make a function called pmap_remove takes m, key returns m2
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] != key then
let out = list_push(out, m[i])
end
let i = i + 1
end
return out
end
make a function called pmap_keys takes m returns keys
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
let out = list_push(out, m[i][0])
let i = i + 1
end
return out
end
make a function called pmap_size takes m returns n
return to_num(list_len(m))
end
# =============================================================================
# schema_bdd: attach a Z-notation-style schema (declared state + a general
# invariant, plus named operations each with a require/ensure precondition/
# postcondition) to a BDD feature, and check a concrete Given/When/Then
# scenario as a WITNESS of that schema -- catching a scenario whose own
# Given already violates an operation's precondition, before any of the
# scenario's own Then assertions even run.
#
# Design decisions, and why (see the implementation plan this file was
# built from for the full investigation):
#
# - A schema's invariant/require/ensure are ORDINARY, separately-compiled
# PatLang functions returning bool, referenced by name string and called
# via apply() -- never the literal `require`/`ensure` keywords. Those
# keywords lower to contract_check (rust-runtime/src/ir/hosts.rs), which
# is FATAL on violation (confirmed directly, and independently documented
# in self_hosting/lib/primitive_registry.patlang's own header) -- unusable
# for a check that must report a diagnosis and keep running, the same
# reason primitive_registry.patlang's own try_ wrappers never use
# require/ensure for their real failure path either.
#
# - apply() has no argument-spread form (confirmed in
# rust-runtime/src/ir/interpreter.rs): a function written once, generic
# over any schema's own number of state variables, cannot pass one
# positional argument per variable. State and inputs are therefore always
# bundled as a single list argument: invariant_fn(state_list),
# require_fn(state_list, input_list), ensure_fn(before_list, after_list,
# input_list).
#
# - Binding extends the existing set_var/get("__vars", ...) convention
# every Gherkin step in this codebase already uses (there is no
# parametrized step-matching mechanism anywhere to extend instead --
# step() dispatch is exact-literal-text, zero-argument, confirmed across
# every existing _bdd_demo.patlang file). A scenario's own Given/Then step
# functions call schema_bind_state/schema_bind_input directly.
#
# HARD RULE: self_hosting/lib/test.patlang's run_feature never resets bound
# vars between scenarios (only t_skipping/t_pending_tags are per-scenario).
# Every scenario's Given/Then must explicitly rebind every declared state
# variable and input itself, every time -- including restating an unchanged
# variable in Then. Do not rely on a value surviving from a previous
# scenario, and do not assume an unmentioned variable defaults to its
# "before" value in Then -- an omitted rebind reads back as the not-yet-
# bound sentinel (see schema_state_values below), which will correctly
# fail the check rather than silently pass, but the failure will look like
# a real violation unless this rule is followed.
# =============================================================================
# Single, process-wide registry (like Step's own new("Step", text)
# convention -- schemas and operations are inherently global, not
# per-instance, so one lazily-created Dict is simpler than
# primitive_registry.patlang's counter-named multi-registry support, which
# this doesn't need).
make a function called schema_registry returns registry
let existing = get("__vars", "schema_bdd_registry_obj")
if existing then
return existing
end
let registry = new("Dict", "schema_bdd_registry")
set_var("schema_bdd_registry_obj", registry)
return registry
end
# state_var_names: list of strings naming the schema's declared state.
# invariant_fn: name of a function taking ONE argument (the state values,
# in the same order as state_var_names) and returning bool.
make a function called schema_define takes name, state_var_names, invariant_fn returns done
send(schema_registry(), "set", name + "__schema", [state_var_names, invariant_fn])
return true
end
make a function called schema_lookup takes name returns entry
return get(schema_registry(), name + "__schema")
end
# input_names: list of strings naming the operation's declared inputs.
# require_fn: name of a function taking (state_values, input_values),
# returning bool -- the operation's precondition.
# ensure_fn: name of a function taking (before_values, after_values,
# input_values), returning bool -- the operation's postcondition.
make a function called schema_operation takes schema_name, op_name, input_names, require_fn, ensure_fn returns done
send(schema_registry(), "set", schema_name + "::" + op_name + "__op", [schema_name, input_names, require_fn, ensure_fn])
return true
end
make a function called schema_lookup_operation takes schema_name, op_name returns entry
return get(schema_registry(), schema_name + "::" + op_name + "__op")
end
# ---- binding: a scenario's own step functions call these ----
make a function called schema_bind_state takes schema, var_name, phase, value returns done
set_var(schema + "__" + var_name + "__" + phase, value)
return true
end
make a function called schema_bind_input takes schema, op, param_name, value returns done
set_var(schema + "__" + op + "__in__" + param_name, value)
return true
end
# Not-yet-bound sentinel: get("__vars", ...) on an unset key -- same
# absence convention already relied on throughout this codebase (e.g.
# gherkin_contracts.patlang's `already_bound` check). Never distinguishes
# "never bound" from "bound to this same falsy value"; low risk for the
# LibraryLoans-shaped schemas this was designed against (string/list-typed
# state), a documented hard limit for anything boolean- or zero-valued.
make a function called schema_state_values takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__before"))
let i = i + 1
end
return out
end
make a function called schema_state_values_after takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__after"))
let i = i + 1
end
return out
end
make a function called schema_input_values takes schema, op, input_names returns values
let out = []
let i = 0
let n = to_num(list_len(input_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + op + "__in__" + input_names[i]))
let i = i + 1
end
return out
end
# ---- the witness check ----
#
# Four-stage check, each stage its own diagnosis tag, checked in the order
# a Z schema's own reasoning goes: is the state even valid to start from,
# does this operation's precondition actually hold for it, does the
# claimed resulting state stay valid, and does the operation's own
# postcondition connect before to after correctly. Returns [tag, payload],
# the same 2-element diagnosis shape as synth5_induce
# (self_hosting/lib/synthesis_lgg.patlang) -- deliberately, so
# schema_format_diagnosis below can mirror synth5_format_diagnosis's own
# real structure rather than inventing a new diagnosis style.
# The value-based core: takes state_before/inputs/state_after directly
# rather than reading them out of the __vars binding store. schema_check
# below is the BDD-scenario-shaped wrapper around this; this function is
# what any OTHER caller (a synthesis harness, a future GOAP or ILP hook --
# see the implementation plan's Phase 3) should call directly instead of
# faking a scenario's Given/When/Then bindings just to reach schema_check.
make a function called schema_check_values takes schema_name, op_name, state_before, inputs, state_after returns diagnosis
let schema_entry = schema_lookup(schema_name)
let invariant_fn = schema_entry[1]
let op_entry = schema_lookup_operation(schema_name, op_name)
let require_fn = op_entry[2]
let ensure_fn = op_entry[3]
if apply(invariant_fn, state_before) == false then
return ["invariant_violated_before", [schema_name, state_before]]
end
if apply(require_fn, state_before, inputs) == false then
return ["precondition_violated", [op_name, state_before, inputs]]
end
if apply(invariant_fn, state_after) == false then
return ["invariant_violated_after", [schema_name, state_after]]
end
if apply(ensure_fn, state_before, state_after, inputs) == false then
return ["postcondition_violated", [op_name, state_before, state_after, inputs]]
end
return ["ok", [op_name, state_before, state_after, inputs]]
end
make a function called schema_check takes schema_name, op_name returns diagnosis
let schema_entry = schema_lookup(schema_name)
let state_var_names = schema_entry[0]
let op_entry = schema_lookup_operation(schema_name, op_name)
let input_names = op_entry[1]
let state_before = schema_state_values(schema_name, state_var_names)
let inputs = schema_input_values(schema_name, op_name, input_names)
let state_after = schema_state_values_after(schema_name, state_var_names)
return schema_check_values(schema_name, op_name, state_before, inputs, state_after)
end
# Mirrors synth5_format_diagnosis's real structure (synthesis_lgg.patlang):
# switch on diagnosis[0], build a human-readable question per case. Returns
# a plain list of question strings (empty for "ok"), same shape as
# synth5_induce_and_ask's own return convention.
make a function called schema_format_diagnosis takes schema_name, op_name, diagnosis returns questions
let tag = diagnosis[0]
let payload = diagnosis[1]
if tag == "invariant_violated_before" then
return ["This scenario's Given already leaves " + schema_name + " in a state that violates its own invariant, before " + op_name + " is even checked. Is the Given wrong, or is the invariant too strict?"]
end
if tag == "precondition_violated" then
return ["Scenario claims " + op_name + " can run from this Given, but " + op_name + "'s own precondition returned false for these inputs. Is the Given wrong, or is " + op_name + "'s precondition too strict -- or is this scenario meant to test a rejection path, which needs its own operation schema rather than " + op_name + "'s?"]
end
if tag == "invariant_violated_after" then
return ["Applying " + op_name + " produces a state that violates " + schema_name + "'s invariant. Is the Then clause's claimed resulting state wrong, or does " + op_name + " need a stronger precondition to rule this case out?"]
end
if tag == "postcondition_violated" then
return ["The before/after state this scenario claims for " + op_name + " does not satisfy its own postcondition. Is the Then clause's claimed resulting state wrong, or is " + op_name + "'s postcondition wrong?"]
end
return []
end
# ---- synthesis integration (implementation plan Phase 3) ----
#
# Opt-in registration linking a SYNTHESIZED function's name to a schema
# operation it's meant to satisfy, plus a "harness" function name that
# knows how to actually exercise it: harness_fn takes (func_name) and
# returns [state_before, inputs, state_after] by calling the synthesized
# function itself (via apply) on some concrete test input and observing
# what it does. This is necessarily domain-specific -- there is no way to
# derive it generically -- so it must be supplied by whoever registers the
# hook, not synthesized here.
#
# Deliberately separate from schema_operation itself: a schema operation
# describes the CONTRACT; a synthesis hook additionally says "and here's
# how to actually run a candidate implementation against it," which only
# matters once there's a synthesized candidate to check, not for the
# scenario-witnessing use in schema_check above.
make a function called schema_register_synthesis_check takes func_name, schema_name, op_name, harness_fn returns done
send(schema_registry(), "set", func_name + "__synthesis_check", [schema_name, op_name, harness_fn])
return true
end
make a function called schema_lookup_synthesis_check takes func_name returns entry
return get(schema_registry(), func_name + "__synthesis_check")
end
# Runs a registered synthesis hook for func_name and returns its
# schema_check_values diagnosis directly. Callers with no hook registered
# for func_name should skip calling this entirely (schema_lookup_
# synthesis_check(func_name) is falsy) rather than call it and inspect the
# result -- there is no "no hook registered" diagnosis tag, since this
# function assumes a hook exists.
make a function called schema_run_synthesis_check takes func_name returns diagnosis
let hook = schema_lookup_synthesis_check(func_name)
let schema_name = hook[0]
let op_name = hook[1]
let harness_fn = hook[2]
let triple = apply(harness_fn, func_name)
return schema_check_values(schema_name, op_name, triple[0], triple[1], triple[2])
end
# Selftest for self_hosting/lib/schema_bdd.patlang: the LibraryLoans/
# BorrowBook worked example from the teaching site's "Schemas and
# Scenarios" page and the implementation plan built from it. Proven through
# the real Gherkin runner (run_feature), not by calling schema_check
# directly -- the point is proving the BDD-integration path works, not
# just the underlying function.
#
# Run from the repo root:
# rust-runtime/target/release/pat --ir-run self_hosting/schema_bdd_selftest.patlang
# ---- LibraryLoans schema: invariant + BorrowBook's precondition/postcondition ----
#
# state = [books, borrowedBy]. Invariant: every title on loan is a title
# the library actually holds -- dom borrowedBy subset_of books, in the
# teaching page's own Z notation.
make a function called schema_inv_library_loans takes state returns ok
let books = state[0]
let borrowedBy = state[1]
let keys = pmap_keys(borrowedBy)
let i = 0
let n = to_num(list_len(keys))
while i < n do
if pset_contains(books, keys[i]) == false then
return false
end
let i = i + 1
end
return true
end
# BorrowBook's precondition: the title exists and isn't already on loan.
make a function called schema_require_borrow_book takes state, inputs returns ok
let books = state[0]
let borrowedBy = state[1]
let title = inputs[0]
if pset_contains(books, title) == false then
return false
end
return pmap_has(borrowedBy, title) == false
end
# BorrowBook's postcondition: the set of titles is unchanged (borrowing
# doesn't add or remove books), and borrowedBy after equals borrowedBy
# before with exactly [title -> member] added.
make a function called schema_ensure_borrow_book takes before, after, inputs returns ok
let books_before = before[0]
let borrowedBy_before = before[1]
let books_after = after[0]
let borrowedBy_after = after[1]
let title = inputs[0]
let member = inputs[1]
if pset_equal(books_before, books_after) == false then
return false
end
let expected_after = pmap_put(borrowedBy_before, title, member)
return schema_pmap_equal(expected_after, borrowedBy_after)
end
# pmap.patlang has no pmap_equal (not needed by pset/pmap's own selftests);
# small enough to define locally here rather than grow pmap.patlang's
# surface for a single caller.
make a function called schema_pmap_equal takes a, b returns eq
if pmap_size(a) != pmap_size(b) then
return false
end
let keys = pmap_keys(a)
let i = 0
let n = to_num(list_len(keys))
while i < n do
let k = keys[i]
if pmap_has(b, k) == false then
return false
end
if pmap_get(a, k) != pmap_get(b, k) then
return false
end
let i = i + 1
end
return true
end
make a function called register_library_loans_schema returns done
schema_define("LibraryLoans", ["books", "borrowedBy"], "schema_inv_library_loans")
schema_operation("LibraryLoans", "BorrowBook", ["title", "member"], "schema_require_borrow_book", "schema_ensure_borrow_book")
return true
end
# ---- step definitions ----
make a function called st_given_library_holds_dune returns done
schema_bind_state("LibraryLoans", "books", "before", pset_add(pset_new(), "Dune"))
return true
end
make a function called st_given_dune_not_borrowed returns done
schema_bind_state("LibraryLoans", "borrowedBy", "before", pmap_new())
return true
end
make a function called st_given_dune_borrowed_by_okonkwo returns done
schema_bind_state("LibraryLoans", "borrowedBy", "before", pmap_put(pmap_new(), "Dune", "S. Okonkwo"))
return true
end
make a function called st_when_diallo_borrows_dune returns done
schema_bind_input("LibraryLoans", "BorrowBook", "title", "Dune")
schema_bind_input("LibraryLoans", "BorrowBook", "member", "A. Diallo")
return true
end
# The scenario's own claimed outcome: books unchanged, Dune now on loan to
# Diallo. Restates books explicitly per schema_bdd.patlang's hard rule
# (every scenario must rebind every declared state var in Then, never rely
# on a default).
make a function called st_then_dune_borrowed_by_diallo returns done
let before_books = get("__vars", "LibraryLoans__books__before")
schema_bind_state("LibraryLoans", "books", "after", before_books)
let before_borrowedBy = get("__vars", "LibraryLoans__borrowedBy__before")
schema_bind_state("LibraryLoans", "borrowedBy", "after", pmap_put(before_borrowedBy, "Dune", "A. Diallo"))
let diagnosis = schema_check("LibraryLoans", "BorrowBook")
check("borrowing an available title is accepted by the schema", diagnosis[0], "ok")
return true
end
# Deliberately does NOT bind an "after" state -- schema_check's own
# ordering (invariant -> precondition -> invariant-after -> postcondition)
# means a precondition failure is reported before "after" is ever read, so
# there is nothing to claim here: the request never gets that far.
make a function called st_then_system_rejects_request returns done
let diagnosis = schema_check("LibraryLoans", "BorrowBook")
check("borrowing an already-borrowed title is rejected by the schema", diagnosis[0], "precondition_violated")
let questions = schema_format_diagnosis("LibraryLoans", "BorrowBook", diagnosis)
check("a rejected scenario gets exactly one diagnostic question", to_num(list_len(questions)), 1)
check("the diagnostic question is non-empty", questions[0].length > 0, true)
return true
end
make a function called register_library_loans_steps returns done
step("the library holds \"Dune\"", "st_given_library_holds_dune")
step("\"Dune\" is not currently borrowed", "st_given_dune_not_borrowed")
step("\"Dune\" is currently borrowed by \"S. Okonkwo\"", "st_given_dune_borrowed_by_okonkwo")
step("\"A. Diallo\" borrows \"Dune\"", "st_when_diallo_borrows_dune")
step("\"Dune\" is borrowed by \"A. Diallo\"", "st_then_dune_borrowed_by_diallo")
step("the system rejects the request", "st_then_system_rejects_request")
return true
end
make a function called library_loans_feature returns text
return "Feature: Library loans
Scenario: Borrowing an available book
Given the library holds \"Dune\"
And \"Dune\" is not currently borrowed
When \"A. Diallo\" borrows \"Dune\"
Then \"Dune\" is borrowed by \"A. Diallo\"
Scenario: Borrowing an already-borrowed book
Given the library holds \"Dune\"
And \"Dune\" is currently borrowed by \"S. Okonkwo\"
When \"A. Diallo\" borrows \"Dune\"
Then the system rejects the request
"
end
make a function called run_schema_bdd_selftest returns ok
t_init()
register_library_loans_schema()
register_library_loans_steps()
run_feature(library_loans_feature())
t_report()
return get("__vars", "t_fail") == 0
end
run_schema_bdd_selftest()
# =============================================================================
# Test framework (Stage 1 dialect): unit assertions plus a Gherkin-style
# feature runner. Step definitions are registered in the object store keyed
# by their text; features are plain text dispatched line by line, so the
# same framework covers unit, integration, and behaviour tests.
# =============================================================================
make a function called t_init returns done
set_var("t_pass", 0)
set_var("t_fail", 0)
set_var("t_tagfilter", "")
set_var("t_pending_tags", "")
set_var("t_skipping", 0)
return true
end
make a function called contains_text takes hay, needle returns r
if needle.length > hay.length then
return false
end
let i = 0
while i <= hay.length - needle.length do
if substr(hay, i, needle.length) == needle then
return true
end
let i = i + 1
end
return false
end
make a function called check takes label, actual, expected returns ok
if actual == expected then
set_var("t_pass", get("__vars", "t_pass") + 1)
print(" ok: " + label)
return true
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: " + label + " (got " + actual + ", want " + expected + ")")
return false
end
end
make a function called t_report returns ok
let p = get("__vars", "t_pass")
let f = get("__vars", "t_fail")
print("tests: " + p + " passed, " + f + " failed")
if f == 0 then
print("ALL TESTS PASSED")
return true
else
print("TESTS FAILED")
return false
end
end
# ---- Gherkin runner ----
# Register a step: step("a fresh till", "st_fresh_till")
make a function called step takes text, fname returns done
new("Step", text)
send(text, "set", "fn", fname)
return true
end
make a function called starts_with takes s, prefix returns r
if s.length < prefix.length then
return false
end
return substr(s, 0, prefix.length) == prefix
end
make a function called trim_left takes s returns out
let i = 0
let scanning = true
while (i < s.length) and scanning do
let c = char_code(s, i)
if (c == 32) or (c == 9) then
let i = i + 1
else
let scanning = false
end
end
return substr(s, i, s.length - i)
end
# Strip a Gherkin keyword; returns the step text or "" if not a step line
make a function called step_text takes line returns out
if starts_with(line, "Given ") then
return substr(line, 6, line.length - 6)
end
if starts_with(line, "When ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "Then ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "And ") then
return substr(line, 4, line.length - 4)
end
return ""
end
# Run only scenarios whose preceding @tag line contains `tag` ("" = all)
make a function called run_feature_tagged takes feature, tag returns ok
set_var("t_tagfilter", tag)
return run_feature(feature)
end
make a function called run_feature_file takes path returns ok
return run_feature(read_file(path))
end
make a function called run_feature takes feature returns ok
let h = str_intern(feature)
let n = sc_len(h)
let i = 0
let line = sb_new()
while i <= n do
let c = sc_code(h, i)
if (c == 10) or (c == -1) then
let raw = trim_left(sb_str(line))
let line = sb_new()
if starts_with(raw, "@") then
set_var("t_pending_tags", raw)
end
if starts_with(raw, "Feature:") then
print(raw)
end
if starts_with(raw, "Scenario:") then
let filter = get("__vars", "t_tagfilter")
let tags = get("__vars", "t_pending_tags")
set_var("t_pending_tags", "")
if filter then
if tags then
if contains_text(tags, filter) then
set_var("t_skipping", 0)
print(raw + " [" + tags + "]")
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 0)
print(raw)
end
else
let text = step_text(raw)
if (text != "") and (get("__vars", "t_skipping") != 1) then
if starts_with(text, "require ") then
handle_contract_step("require", substr(text, 8, text.length - 8))
else
if starts_with(text, "ensure ") then
handle_contract_step("ensure", substr(text, 7, text.length - 7))
else
let fname = get(text, "fn")
if fname then
apply(fname)
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: undefined step: " + text)
end
end
end
end
end
let i = i + 1
else
if c == 13 then
let i = i + 1
else
sb_push(line, sc_char(h, i))
let i = i + 1
end
end
end
return true
end
# =============================================================================
# pset: a minimal Set<T> built on a plain list with linear-scan membership,
# the same idiom already used by depgraph_list_contains and
# synth5_list_contains (self_hosting/lib/depgraph.patlang,
# self_hosting/lib/synthesis_lgg.patlang) rather than a new primitive --
# PatLang has no native Set type.
#
# Functional / return-new-collection style throughout (mirrors list_push's
# own semantics): every mutating-sounding operation returns a NEW list
# rather than mutating in place. This matters specifically for schema_bdd's
# before/after snapshotting, which needs an independent copy of state at
# two points in time -- an in-place structure would make "before" silently
# turn into "after" once the operation runs.
# =============================================================================
make a function called pset_new returns s
return []
end
make a function called pset_contains takes s, item returns found
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] == item then
return true
end
let i = i + 1
end
return false
end
make a function called pset_add takes s, item returns s2
if pset_contains(s, item) then
return s
end
return list_push(s, item)
end
make a function called pset_remove takes s, item returns s2
let out = []
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] != item then
let out = list_push(out, s[i])
end
let i = i + 1
end
return out
end
make a function called pset_size takes s returns n
return to_num(list_len(s))
end
# Identity -- exposed so callers iterating a set's members don't need to
# know it's "just a list" internally; the name documents intent at the
# call site instead.
make a function called pset_to_list takes s returns items
return s
end
# Set equality: same size, and every member of a is in b. Given pset_add's
# own dedup, equal size plus one-directional containment implies the
# reverse containment too (no way for b to hold something a doesn't
# without differing in size).
make a function called pset_equal takes a, b returns eq
if pset_size(a) != pset_size(b) then
return false
end
let i = 0
let n = to_num(list_len(a))
while i < n do
if pset_contains(b, a[i]) == false then
return false
end
let i = i + 1
end
return true
end
# =============================================================================
# pmap: a minimal Map<K,V> as an association list of [key, value] pairs with
# linear-scan lookup -- the exact same idiom already proven in
# self_hosting/lib/depgraph.patlang's depgraph_map_get/depgraph_map_append
# (a file->list-of-files map), generalised here to an arbitrary value type
# and given a full put/get/has/remove/keys surface. PatLang has no native
# Map/Dict type usable from self-hosted code at this scale.
#
# Functional / return-new-collection style throughout, for the same reason
# as pset.patlang: schema_bdd needs an independent before/after snapshot of
# state, which an in-place-mutating map wouldn't give for free.
#
# pmap_get returns [] (an empty list) for a missing key -- the same
# not-found sentinel depgraph_map_get already uses, not a distinct "Unit"
# literal (PatLang's dialect has no source-level Unit/nil literal to write
# directly). This is INHERENTLY AMBIGUOUS if a real stored value could
# itself be an empty list: pmap_has is the only call that actually
# distinguishes "absent" from "present but happens to look like the
# sentinel" -- never infer presence from pmap_get's return value alone.
# =============================================================================
make a function called pmap_new returns m
return []
end
make a function called pmap_has takes m, key returns found
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return true
end
let i = i + 1
end
return false
end
# See the module-level warning above: [] means "not found" here, which is
# indistinguishable from a genuinely stored empty-list value. Call
# pmap_has first whenever that distinction matters.
make a function called pmap_get takes m, key returns value
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return m[i][1]
end
let i = i + 1
end
return []
end
make a function called pmap_put takes m, key, value returns m2
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return list_set(m, i, [key, value])
end
let i = i + 1
end
return list_push(m, [key, value])
end
make a function called pmap_remove takes m, key returns m2
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] != key then
let out = list_push(out, m[i])
end
let i = i + 1
end
return out
end
make a function called pmap_keys takes m returns keys
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
let out = list_push(out, m[i][0])
let i = i + 1
end
return out
end
make a function called pmap_size takes m returns n
return to_num(list_len(m))
end
# =============================================================================
# schema_bdd: attach a Z-notation-style schema (declared state + a general
# invariant, plus named operations each with a require/ensure precondition/
# postcondition) to a BDD feature, and check a concrete Given/When/Then
# scenario as a WITNESS of that schema -- catching a scenario whose own
# Given already violates an operation's precondition, before any of the
# scenario's own Then assertions even run.
#
# Design decisions, and why (see the implementation plan this file was
# built from for the full investigation):
#
# - A schema's invariant/require/ensure are ORDINARY, separately-compiled
# PatLang functions returning bool, referenced by name string and called
# via apply() -- never the literal `require`/`ensure` keywords. Those
# keywords lower to contract_check (rust-runtime/src/ir/hosts.rs), which
# is FATAL on violation (confirmed directly, and independently documented
# in self_hosting/lib/primitive_registry.patlang's own header) -- unusable
# for a check that must report a diagnosis and keep running, the same
# reason primitive_registry.patlang's own try_ wrappers never use
# require/ensure for their real failure path either.
#
# - apply() has no argument-spread form (confirmed in
# rust-runtime/src/ir/interpreter.rs): a function written once, generic
# over any schema's own number of state variables, cannot pass one
# positional argument per variable. State and inputs are therefore always
# bundled as a single list argument: invariant_fn(state_list),
# require_fn(state_list, input_list), ensure_fn(before_list, after_list,
# input_list).
#
# - Binding extends the existing set_var/get("__vars", ...) convention
# every Gherkin step in this codebase already uses (there is no
# parametrized step-matching mechanism anywhere to extend instead --
# step() dispatch is exact-literal-text, zero-argument, confirmed across
# every existing _bdd_demo.patlang file). A scenario's own Given/Then step
# functions call schema_bind_state/schema_bind_input directly.
#
# HARD RULE: self_hosting/lib/test.patlang's run_feature never resets bound
# vars between scenarios (only t_skipping/t_pending_tags are per-scenario).
# Every scenario's Given/Then must explicitly rebind every declared state
# variable and input itself, every time -- including restating an unchanged
# variable in Then. Do not rely on a value surviving from a previous
# scenario, and do not assume an unmentioned variable defaults to its
# "before" value in Then -- an omitted rebind reads back as the not-yet-
# bound sentinel (see schema_state_values below), which will correctly
# fail the check rather than silently pass, but the failure will look like
# a real violation unless this rule is followed.
# =============================================================================
# Single, process-wide registry (like Step's own new("Step", text)
# convention -- schemas and operations are inherently global, not
# per-instance, so one lazily-created Dict is simpler than
# primitive_registry.patlang's counter-named multi-registry support, which
# this doesn't need).
make a function called schema_registry returns registry
let existing = get("__vars", "schema_bdd_registry_obj")
if existing then
return existing
end
let registry = new("Dict", "schema_bdd_registry")
set_var("schema_bdd_registry_obj", registry)
return registry
end
# state_var_names: list of strings naming the schema's declared state.
# invariant_fn: name of a function taking ONE argument (the state values,
# in the same order as state_var_names) and returning bool.
make a function called schema_define takes name, state_var_names, invariant_fn returns done
send(schema_registry(), "set", name + "__schema", [state_var_names, invariant_fn])
return true
end
make a function called schema_lookup takes name returns entry
return get(schema_registry(), name + "__schema")
end
# input_names: list of strings naming the operation's declared inputs.
# require_fn: name of a function taking (state_values, input_values),
# returning bool -- the operation's precondition.
# ensure_fn: name of a function taking (before_values, after_values,
# input_values), returning bool -- the operation's postcondition.
make a function called schema_operation takes schema_name, op_name, input_names, require_fn, ensure_fn returns done
send(schema_registry(), "set", schema_name + "::" + op_name + "__op", [schema_name, input_names, require_fn, ensure_fn])
return true
end
make a function called schema_lookup_operation takes schema_name, op_name returns entry
return get(schema_registry(), schema_name + "::" + op_name + "__op")
end
# ---- binding: a scenario's own step functions call these ----
make a function called schema_bind_state takes schema, var_name, phase, value returns done
set_var(schema + "__" + var_name + "__" + phase, value)
return true
end
make a function called schema_bind_input takes schema, op, param_name, value returns done
set_var(schema + "__" + op + "__in__" + param_name, value)
return true
end
# Not-yet-bound sentinel: get("__vars", ...) on an unset key -- same
# absence convention already relied on throughout this codebase (e.g.
# gherkin_contracts.patlang's `already_bound` check). Never distinguishes
# "never bound" from "bound to this same falsy value"; low risk for the
# LibraryLoans-shaped schemas this was designed against (string/list-typed
# state), a documented hard limit for anything boolean- or zero-valued.
make a function called schema_state_values takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__before"))
let i = i + 1
end
return out
end
make a function called schema_state_values_after takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__after"))
let i = i + 1
end
return out
end
make a function called schema_input_values takes schema, op, input_names returns values
let out = []
let i = 0
let n = to_num(list_len(input_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + op + "__in__" + input_names[i]))
let i = i + 1
end
return out
end
# ---- the witness check ----
#
# Four-stage check, each stage its own diagnosis tag, checked in the order
# a Z schema's own reasoning goes: is the state even valid to start from,
# does this operation's precondition actually hold for it, does the
# claimed resulting state stay valid, and does the operation's own
# postcondition connect before to after correctly. Returns [tag, payload],
# the same 2-element diagnosis shape as synth5_induce
# (self_hosting/lib/synthesis_lgg.patlang) -- deliberately, so
# schema_format_diagnosis below can mirror synth5_format_diagnosis's own
# real structure rather than inventing a new diagnosis style.
# The value-based core: takes state_before/inputs/state_after directly
# rather than reading them out of the __vars binding store. schema_check
# below is the BDD-scenario-shaped wrapper around this; this function is
# what any OTHER caller (a synthesis harness, a future GOAP or ILP hook --
# see the implementation plan's Phase 3) should call directly instead of
# faking a scenario's Given/When/Then bindings just to reach schema_check.
make a function called schema_check_values takes schema_name, op_name, state_before, inputs, state_after returns diagnosis
let schema_entry = schema_lookup(schema_name)
let invariant_fn = schema_entry[1]
let op_entry = schema_lookup_operation(schema_name, op_name)
let require_fn = op_entry[2]
let ensure_fn = op_entry[3]
if apply(invariant_fn, state_before) == false then
return ["invariant_violated_before", [schema_name, state_before]]
end
if apply(require_fn, state_before, inputs) == false then
return ["precondition_violated", [op_name, state_before, inputs]]
end
if apply(invariant_fn, state_after) == false then
return ["invariant_violated_after", [schema_name, state_after]]
end
if apply(ensure_fn, state_before, state_after, inputs) == false then
return ["postcondition_violated", [op_name, state_before, state_after, inputs]]
end
return ["ok", [op_name, state_before, state_after, inputs]]
end
make a function called schema_check takes schema_name, op_name returns diagnosis
let schema_entry = schema_lookup(schema_name)
let state_var_names = schema_entry[0]
let op_entry = schema_lookup_operation(schema_name, op_name)
let input_names = op_entry[1]
let state_before = schema_state_values(schema_name, state_var_names)
let inputs = schema_input_values(schema_name, op_name, input_names)
let state_after = schema_state_values_after(schema_name, state_var_names)
return schema_check_values(schema_name, op_name, state_before, inputs, state_after)
end
# Mirrors synth5_format_diagnosis's real structure (synthesis_lgg.patlang):
# switch on diagnosis[0], build a human-readable question per case. Returns
# a plain list of question strings (empty for "ok"), same shape as
# synth5_induce_and_ask's own return convention.
make a function called schema_format_diagnosis takes schema_name, op_name, diagnosis returns questions
let tag = diagnosis[0]
let payload = diagnosis[1]
if tag == "invariant_violated_before" then
return ["This scenario's Given already leaves " + schema_name + " in a state that violates its own invariant, before " + op_name + " is even checked. Is the Given wrong, or is the invariant too strict?"]
end
if tag == "precondition_violated" then
return ["Scenario claims " + op_name + " can run from this Given, but " + op_name + "'s own precondition returned false for these inputs. Is the Given wrong, or is " + op_name + "'s precondition too strict -- or is this scenario meant to test a rejection path, which needs its own operation schema rather than " + op_name + "'s?"]
end
if tag == "invariant_violated_after" then
return ["Applying " + op_name + " produces a state that violates " + schema_name + "'s invariant. Is the Then clause's claimed resulting state wrong, or does " + op_name + " need a stronger precondition to rule this case out?"]
end
if tag == "postcondition_violated" then
return ["The before/after state this scenario claims for " + op_name + " does not satisfy its own postcondition. Is the Then clause's claimed resulting state wrong, or is " + op_name + "'s postcondition wrong?"]
end
return []
end
# ---- synthesis integration (implementation plan Phase 3) ----
#
# Opt-in registration linking a SYNTHESIZED function's name to a schema
# operation it's meant to satisfy, plus a "harness" function name that
# knows how to actually exercise it: harness_fn takes (func_name) and
# returns [state_before, inputs, state_after] by calling the synthesized
# function itself (via apply) on some concrete test input and observing
# what it does. This is necessarily domain-specific -- there is no way to
# derive it generically -- so it must be supplied by whoever registers the
# hook, not synthesized here.
#
# Deliberately separate from schema_operation itself: a schema operation
# describes the CONTRACT; a synthesis hook additionally says "and here's
# how to actually run a candidate implementation against it," which only
# matters once there's a synthesized candidate to check, not for the
# scenario-witnessing use in schema_check above.
make a function called schema_register_synthesis_check takes func_name, schema_name, op_name, harness_fn returns done
send(schema_registry(), "set", func_name + "__synthesis_check", [schema_name, op_name, harness_fn])
return true
end
make a function called schema_lookup_synthesis_check takes func_name returns entry
return get(schema_registry(), func_name + "__synthesis_check")
end
# Runs a registered synthesis hook for func_name and returns its
# schema_check_values diagnosis directly. Callers with no hook registered
# for func_name should skip calling this entirely (schema_lookup_
# synthesis_check(func_name) is falsy) rather than call it and inspect the
# result -- there is no "no hook registered" diagnosis tag, since this
# function assumes a hook exists.
make a function called schema_run_synthesis_check takes func_name returns diagnosis
let hook = schema_lookup_synthesis_check(func_name)
let schema_name = hook[0]
let op_name = hook[1]
let harness_fn = hook[2]
let triple = apply(harness_fn, func_name)
return schema_check_values(schema_name, op_name, triple[0], triple[1], triple[2])
end
# =============================================================================
# LibraryLoans, fully worked: a richer Z/BDD hybrid schema than the plan's
# own minimal selftest (self_hosting/schema_bdd_selftest.patlang) --
# three pieces of interacting state, a multi-clause precondition, a
# cross-cutting invariant (no member may exceed three simultaneous
# loans), and enough scenarios to exercise four of the five outcomes
# schema_check can return: invariant_violated_before, precondition_
# violated, postcondition_violated, and ok. The fifth,
# invariant_violated_after, is exercised in schema_bdd_selftest.patlang. Companion to the parslow.net
# "Schemas and Scenarios" page.
#
# State: books (Set<Title>), borrowedBy (Map<Title,Member>), loanCount
# (Map<Member,Int>) -- loanCount is tracked SEPARATELY from borrowedBy
# rather than recomputed from it on every check, the same design choice
# a real system would make (a denormalised, maintained count rather than
# a full table scan every time), which is exactly why BorrowBook/
# ReturnBook's own postconditions -- not just the invariant -- have real
# work to do keeping the two consistent with each other.
#
# Run from the repo root:
# rust-runtime/target/release/pat --ir-run self_hosting/examples/library_loans_schema_demo.patlang
# ---- LibraryLoans: invariant + BorrowBook/ReturnBook contracts ----
make a function called ll_loan_count_of takes loan_count, member returns n
if pmap_has(loan_count, member) then
return pmap_get(loan_count, member)
end
return 0
end
# Invariant: every title on loan is one the library actually holds, and
# no member exceeds three simultaneous loans -- the second half is a
# genuine cross-cutting business rule, not derivable from the first.
make a function called ll_invariant takes state returns ok
let books = state[0]
let borrowed_by = state[1]
let loan_count = state[2]
let keys = pmap_keys(borrowed_by)
let i = 0
let n = to_num(list_len(keys))
while i < n do
if pset_contains(books, keys[i]) == false then
return false
end
let i = i + 1
end
let members = pmap_keys(loan_count)
let j = 0
let mn = to_num(list_len(members))
while j < mn do
let count = pmap_get(loan_count, members[j])
if (count < 0) or (count > 3) then
return false
end
let j = j + 1
end
return true
end
# BorrowBook's precondition is three independent clauses, checked
# together -- exactly the shape a hand-written scenario can violate in
# three genuinely different ways, each needing its own scenario to
# demonstrate (see the feature text below).
make a function called ll_require_borrow takes state, inputs returns ok
let books = state[0]
let borrowed_by = state[1]
let loan_count = state[2]
let title = inputs[0]
let member = inputs[1]
if pset_contains(books, title) == false then
return false
end
if pmap_has(borrowed_by, title) then
return false
end
return ll_loan_count_of(loan_count, member) < 3
end
make a function called ll_ensure_borrow takes before, after, inputs returns ok
let books_before = before[0]
let borrowed_by_before = before[1]
let loan_count_before = before[2]
let books_after = after[0]
let borrowed_by_after = after[1]
let loan_count_after = after[2]
let title = inputs[0]
let member = inputs[1]
if pset_equal(books_before, books_after) == false then
return false
end
let expected_borrowed_by = pmap_put(borrowed_by_before, title, member)
if schema_pmap_equal(expected_borrowed_by, borrowed_by_after) == false then
return false
end
let expected_count = ll_loan_count_of(loan_count_before, member) + 1
let expected_loan_count = pmap_put(loan_count_before, member, expected_count)
return schema_pmap_equal(expected_loan_count, loan_count_after)
end
# ReturnBook takes only the title -- the member is recovered from
# borrowedBy, exactly the way a real return desk works (you hand back
# the book, not a signed claim of who you are).
make a function called ll_require_return takes state, inputs returns ok
let borrowed_by = state[1]
let title = inputs[0]
return pmap_has(borrowed_by, title)
end
make a function called ll_ensure_return takes before, after, inputs returns ok
let books_before = before[0]
let borrowed_by_before = before[1]
let loan_count_before = before[2]
let books_after = after[0]
let borrowed_by_after = after[1]
let loan_count_after = after[2]
let title = inputs[0]
let member = pmap_get(borrowed_by_before, title)
if pset_equal(books_before, books_after) == false then
return false
end
let expected_borrowed_by = pmap_remove(borrowed_by_before, title)
if schema_pmap_equal(expected_borrowed_by, borrowed_by_after) == false then
return false
end
let expected_count = ll_loan_count_of(loan_count_before, member) - 1
let expected_loan_count = pmap_put(loan_count_before, member, expected_count)
return schema_pmap_equal(expected_loan_count, loan_count_after)
end
# pmap.patlang has no pmap_equal (see schema_bdd_selftest.patlang's own
# note on this); defined once here, shared by both ensure functions.
make a function called schema_pmap_equal takes a, b returns eq
if pmap_size(a) != pmap_size(b) then
return false
end
let keys = pmap_keys(a)
let i = 0
let n = to_num(list_len(keys))
while i < n do
let k = keys[i]
if pmap_has(b, k) == false then
return false
end
if pmap_get(a, k) != pmap_get(b, k) then
return false
end
let i = i + 1
end
return true
end
make a function called register_library_loans_schema returns done
schema_define("LibraryLoans", ["books", "borrowedBy", "loanCount"], "ll_invariant")
schema_operation("LibraryLoans", "BorrowBook", ["title", "member"], "ll_require_borrow", "ll_ensure_borrow")
schema_operation("LibraryLoans", "ReturnBook", ["title"], "ll_require_return", "ll_ensure_return")
return true
end
# ---- step definitions ----
make a function called ll_standard_catalogue returns s
let s = pset_add(pset_new(), "Dune")
let s = pset_add(s, "Neuromancer")
let s = pset_add(s, "Foundation")
return s
end
make a function called st_given_standard_catalogue returns done
schema_bind_state("LibraryLoans", "books", "before", ll_standard_catalogue())
return true
end
make a function called st_given_no_loans_out returns done
schema_bind_state("LibraryLoans", "borrowedBy", "before", pmap_new())
schema_bind_state("LibraryLoans", "loanCount", "before", pmap_new())
return true
end
make a function called st_given_dune_borrowed_by_okonkwo returns done
schema_bind_state("LibraryLoans", "borrowedBy", "before", pmap_put(pmap_new(), "Dune", "S. Okonkwo"))
schema_bind_state("LibraryLoans", "loanCount", "before", pmap_put(pmap_new(), "S. Okonkwo", 1))
return true
end
# Deliberately consistent, valid "already at the limit" state -- three
# real titles, all correctly attributed to the same member, matching
# what the invariant actually checks (the count, not the specific books).
make a function called st_given_diallo_at_loan_limit returns done
let borrowed_by = pmap_put(pmap_new(), "Dune", "A. Diallo")
let borrowed_by = pmap_put(borrowed_by, "Neuromancer", "A. Diallo")
let borrowed_by = pmap_put(borrowed_by, "Foundation", "A. Diallo")
schema_bind_state("LibraryLoans", "borrowedBy", "before", borrowed_by)
schema_bind_state("LibraryLoans", "loanCount", "before", pmap_put(pmap_new(), "A. Diallo", 3))
return true
end
# Deliberately INVALID state: the record already claims four loans,
# violating the invariant before any operation is even attempted.
make a function called st_given_corrupted_loan_record returns done
schema_bind_state("LibraryLoans", "borrowedBy", "before", pmap_new())
schema_bind_state("LibraryLoans", "loanCount", "before", pmap_put(pmap_new(), "A. Diallo", 4))
return true
end
make a function called st_when_diallo_borrows_foundation returns done
schema_bind_input("LibraryLoans", "BorrowBook", "title", "Foundation")
schema_bind_input("LibraryLoans", "BorrowBook", "member", "A. Diallo")
return true
end
make a function called st_when_diallo_borrows_dune returns done
schema_bind_input("LibraryLoans", "BorrowBook", "title", "Dune")
schema_bind_input("LibraryLoans", "BorrowBook", "member", "A. Diallo")
return true
end
make a function called st_when_okonkwo_returns_dune returns done
schema_bind_input("LibraryLoans", "ReturnBook", "title", "Dune")
return true
end
make a function called st_when_someone_returns_foundation returns done
schema_bind_input("LibraryLoans", "ReturnBook", "title", "Foundation")
return true
end
# ---- Then steps: state the scenario's own claim, then check it ----
make a function called ll_check takes op, expected_tag, label returns done
let diagnosis = schema_check("LibraryLoans", op)
check(label, diagnosis[0], expected_tag)
return true
end
make a function called st_then_foundation_correctly_recorded returns done
let before_books = get("__vars", "LibraryLoans__books__before")
schema_bind_state("LibraryLoans", "books", "after", before_books)
let before_borrowed_by = get("__vars", "LibraryLoans__borrowedBy__before")
schema_bind_state("LibraryLoans", "borrowedBy", "after", pmap_put(before_borrowed_by, "Foundation", "A. Diallo"))
let before_loan_count = get("__vars", "LibraryLoans__loanCount__before")
schema_bind_state("LibraryLoans", "loanCount", "after", pmap_put(before_loan_count, "A. Diallo", ll_loan_count_of(before_loan_count, "A. Diallo") + 1))
ll_check("BorrowBook", "ok", "borrowing an available title, under the loan limit, is accepted")
return true
end
make a function called st_then_borrow_rejected_already_out returns done
ll_check("BorrowBook", "precondition_violated", "borrowing a title someone else already has out is rejected")
return true
end
make a function called st_then_borrow_rejected_at_limit returns done
ll_check("BorrowBook", "precondition_violated", "borrowing while already at the three-loan limit is rejected")
return true
end
make a function called st_then_invariant_flagged_first returns done
ll_check("BorrowBook", "invariant_violated_before", "a corrupted loan record is flagged before the operation is even considered")
return true
end
# Deliberately WRONG claimed outcome: the loan count is claimed to stay
# the same after a successful borrow, which violates BorrowBook's own
# postcondition even though the precondition and invariant are both fine.
make a function called st_then_wrongly_claims_count_unchanged returns done
let before_books = get("__vars", "LibraryLoans__books__before")
schema_bind_state("LibraryLoans", "books", "after", before_books)
let before_borrowed_by = get("__vars", "LibraryLoans__borrowedBy__before")
schema_bind_state("LibraryLoans", "borrowedBy", "after", pmap_put(before_borrowed_by, "Dune", "A. Diallo"))
let before_loan_count = get("__vars", "LibraryLoans__loanCount__before")
schema_bind_state("LibraryLoans", "loanCount", "after", before_loan_count)
ll_check("BorrowBook", "postcondition_violated", "a claimed loan count that doesn't actually increment is rejected")
return true
end
make a function called st_then_return_accepted returns done
let before_books = get("__vars", "LibraryLoans__books__before")
schema_bind_state("LibraryLoans", "books", "after", before_books)
let before_borrowed_by = get("__vars", "LibraryLoans__borrowedBy__before")
schema_bind_state("LibraryLoans", "borrowedBy", "after", pmap_remove(before_borrowed_by, "Dune"))
let before_loan_count = get("__vars", "LibraryLoans__loanCount__before")
schema_bind_state("LibraryLoans", "loanCount", "after", pmap_put(before_loan_count, "S. Okonkwo", ll_loan_count_of(before_loan_count, "S. Okonkwo") - 1))
ll_check("ReturnBook", "ok", "returning a borrowed title is accepted")
return true
end
make a function called st_then_return_rejected_never_out returns done
ll_check("ReturnBook", "precondition_violated", "returning a title nobody borrowed is rejected")
return true
end
make a function called register_library_loans_steps returns done
step("the library holds its standard catalogue", "st_given_standard_catalogue")
step("no member has any books out", "st_given_no_loans_out")
step("\"Dune\" is currently borrowed by \"S. Okonkwo\"", "st_given_dune_borrowed_by_okonkwo")
step("\"A. Diallo\" already has three books out, at the loan limit", "st_given_diallo_at_loan_limit")
step("the loan record incorrectly already shows \"A. Diallo\" with four loans", "st_given_corrupted_loan_record")
step("\"A. Diallo\" borrows \"Foundation\"", "st_when_diallo_borrows_foundation")
step("\"A. Diallo\" borrows \"Dune\"", "st_when_diallo_borrows_dune")
step("\"S. Okonkwo\" returns \"Dune\"", "st_when_okonkwo_returns_dune")
step("someone returns \"Foundation\"", "st_when_someone_returns_foundation")
step("the loan is correctly recorded", "st_then_foundation_correctly_recorded")
step("the borrow is rejected", "st_then_borrow_rejected_already_out")
step("the borrow is rejected for being at the limit", "st_then_borrow_rejected_at_limit")
step("the schema flags the corrupted record, not the borrow attempt", "st_then_invariant_flagged_first")
step("the incorrect record is rejected", "st_then_wrongly_claims_count_unchanged")
step("the return is accepted", "st_then_return_accepted")
step("the return is rejected", "st_then_return_rejected_never_out")
return true
end
make a function called library_loans_feature returns text
return "Feature: Library loans, fully worked
Scenario: Borrowing an available title under the loan limit
Given the library holds its standard catalogue
And no member has any books out
When \"A. Diallo\" borrows \"Foundation\"
Then the loan is correctly recorded
Scenario: Borrowing a title someone else already has out
Given the library holds its standard catalogue
And \"Dune\" is currently borrowed by \"S. Okonkwo\"
When \"A. Diallo\" borrows \"Dune\"
Then the borrow is rejected
Scenario: Borrowing while already at the three-loan limit
Given the library holds its standard catalogue
And \"A. Diallo\" already has three books out, at the loan limit
When \"A. Diallo\" borrows \"Dune\"
Then the borrow is rejected for being at the limit
Scenario: A corrupted loan record is caught before the operation runs
Given the library holds its standard catalogue
And the loan record incorrectly already shows \"A. Diallo\" with four loans
When \"A. Diallo\" borrows \"Foundation\"
Then the schema flags the corrupted record, not the borrow attempt
Scenario: A scenario that under-reports its own effect is caught too
Given the library holds its standard catalogue
And no member has any books out
When \"A. Diallo\" borrows \"Dune\"
Then the incorrect record is rejected
Scenario: Returning a borrowed title
Given the library holds its standard catalogue
And \"Dune\" is currently borrowed by \"S. Okonkwo\"
When \"S. Okonkwo\" returns \"Dune\"
Then the return is accepted
Scenario: Returning a title nobody borrowed
Given the library holds its standard catalogue
And no member has any books out
When someone returns \"Foundation\"
Then the return is rejected
"
end
make a function called run_library_loans_schema_demo returns ok
t_init()
register_library_loans_schema()
register_library_loans_steps()
run_feature(library_loans_feature())
t_report()
return get("__vars", "t_fail") == 0
end
run_library_loans_schema_demo()
# =============================================================================
# Inter-library transfer via GOAP, checked against real resulting state --
# the payoff plan_with_state exists for. The pre-existing gherkin_
# contracts.patlang / goap_verify_contracts system can only check a
# contract against a STRING-PARSED action-label binding, e.g.
# "ship_to_branch(Book=rare_atlas)" -- it has no way to express "and the
# book must not ALSO still be shelved at the origin branch," because that
# needs the plan's full resulting fact set, not one action's own
# parameter. plan_with_state exposes exactly that.
#
# Domain: a rare book moves branch_a -> depot -> branch_b through three
# GOAP actions; a TransferPolicy schema checks the actual resulting
# state, not the plan's own step labels.
#
# Run from the repo root:
# rust-runtime/target/release/pat --ir-run self_hosting/examples/interlibrary_transfer_goap_demo.patlang
# =============================================================================
# Test framework (Stage 1 dialect): unit assertions plus a Gherkin-style
# feature runner. Step definitions are registered in the object store keyed
# by their text; features are plain text dispatched line by line, so the
# same framework covers unit, integration, and behaviour tests.
# =============================================================================
make a function called t_init returns done
set_var("t_pass", 0)
set_var("t_fail", 0)
set_var("t_tagfilter", "")
set_var("t_pending_tags", "")
set_var("t_skipping", 0)
return true
end
make a function called contains_text takes hay, needle returns r
if needle.length > hay.length then
return false
end
let i = 0
while i <= hay.length - needle.length do
if substr(hay, i, needle.length) == needle then
return true
end
let i = i + 1
end
return false
end
make a function called check takes label, actual, expected returns ok
if actual == expected then
set_var("t_pass", get("__vars", "t_pass") + 1)
print(" ok: " + label)
return true
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: " + label + " (got " + actual + ", want " + expected + ")")
return false
end
end
make a function called t_report returns ok
let p = get("__vars", "t_pass")
let f = get("__vars", "t_fail")
print("tests: " + p + " passed, " + f + " failed")
if f == 0 then
print("ALL TESTS PASSED")
return true
else
print("TESTS FAILED")
return false
end
end
# ---- Gherkin runner ----
# Register a step: step("a fresh till", "st_fresh_till")
make a function called step takes text, fname returns done
new("Step", text)
send(text, "set", "fn", fname)
return true
end
make a function called starts_with takes s, prefix returns r
if s.length < prefix.length then
return false
end
return substr(s, 0, prefix.length) == prefix
end
make a function called trim_left takes s returns out
let i = 0
let scanning = true
while (i < s.length) and scanning do
let c = char_code(s, i)
if (c == 32) or (c == 9) then
let i = i + 1
else
let scanning = false
end
end
return substr(s, i, s.length - i)
end
# Strip a Gherkin keyword; returns the step text or "" if not a step line
make a function called step_text takes line returns out
if starts_with(line, "Given ") then
return substr(line, 6, line.length - 6)
end
if starts_with(line, "When ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "Then ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "And ") then
return substr(line, 4, line.length - 4)
end
return ""
end
# Run only scenarios whose preceding @tag line contains `tag` ("" = all)
make a function called run_feature_tagged takes feature, tag returns ok
set_var("t_tagfilter", tag)
return run_feature(feature)
end
make a function called run_feature_file takes path returns ok
return run_feature(read_file(path))
end
make a function called run_feature takes feature returns ok
let h = str_intern(feature)
let n = sc_len(h)
let i = 0
let line = sb_new()
while i <= n do
let c = sc_code(h, i)
if (c == 10) or (c == -1) then
let raw = trim_left(sb_str(line))
let line = sb_new()
if starts_with(raw, "@") then
set_var("t_pending_tags", raw)
end
if starts_with(raw, "Feature:") then
print(raw)
end
if starts_with(raw, "Scenario:") then
let filter = get("__vars", "t_tagfilter")
let tags = get("__vars", "t_pending_tags")
set_var("t_pending_tags", "")
if filter then
if tags then
if contains_text(tags, filter) then
set_var("t_skipping", 0)
print(raw + " [" + tags + "]")
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 0)
print(raw)
end
else
let text = step_text(raw)
if (text != "") and (get("__vars", "t_skipping") != 1) then
if starts_with(text, "require ") then
handle_contract_step("require", substr(text, 8, text.length - 8))
else
if starts_with(text, "ensure ") then
handle_contract_step("ensure", substr(text, 7, text.length - 7))
else
let fname = get(text, "fn")
if fname then
apply(fname)
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: undefined step: " + text)
end
end
end
end
end
let i = i + 1
else
if c == 13 then
let i = i + 1
else
sb_push(line, sc_char(h, i))
let i = i + 1
end
end
end
return true
end
# =============================================================================
# schema_goap_bridge: connects schema_bdd.patlang to the GOAP planner via
# the native plan_with_state host function (rust-runtime/src/ir/hosts.rs)
# -- implementation plan Phase 3, item 3.
#
# goap_verify_contracts (self_hosting/lib/gherkin_contracts.patlang) only
# ever checks a contract clause against STRING-PARSED action-label
# bindings ("scale(X=5)"-style, via gc_extract_binding) -- plan()'s own
# real resulting world-state was always computed internally but discarded
# before reaching PatLang. plan_with_state exposes it directly: [path,
# resulting_facts_list, resulting_fluents_list], resulting_facts_list in
# the same [pred, args_list] shape rule_add's own arguments already use.
# This bridge lets a schema check that ACTUAL resulting state directly,
# instead of re-parsing a label string for a variable binding.
# =============================================================================
# =============================================================================
# schema_bdd: attach a Z-notation-style schema (declared state + a general
# invariant, plus named operations each with a require/ensure precondition/
# postcondition) to a BDD feature, and check a concrete Given/When/Then
# scenario as a WITNESS of that schema -- catching a scenario whose own
# Given already violates an operation's precondition, before any of the
# scenario's own Then assertions even run.
#
# Design decisions, and why (see the implementation plan this file was
# built from for the full investigation):
#
# - A schema's invariant/require/ensure are ORDINARY, separately-compiled
# PatLang functions returning bool, referenced by name string and called
# via apply() -- never the literal `require`/`ensure` keywords. Those
# keywords lower to contract_check (rust-runtime/src/ir/hosts.rs), which
# is FATAL on violation (confirmed directly, and independently documented
# in self_hosting/lib/primitive_registry.patlang's own header) -- unusable
# for a check that must report a diagnosis and keep running, the same
# reason primitive_registry.patlang's own try_ wrappers never use
# require/ensure for their real failure path either.
#
# - apply() has no argument-spread form (confirmed in
# rust-runtime/src/ir/interpreter.rs): a function written once, generic
# over any schema's own number of state variables, cannot pass one
# positional argument per variable. State and inputs are therefore always
# bundled as a single list argument: invariant_fn(state_list),
# require_fn(state_list, input_list), ensure_fn(before_list, after_list,
# input_list).
#
# - Binding extends the existing set_var/get("__vars", ...) convention
# every Gherkin step in this codebase already uses (there is no
# parametrized step-matching mechanism anywhere to extend instead --
# step() dispatch is exact-literal-text, zero-argument, confirmed across
# every existing _bdd_demo.patlang file). A scenario's own Given/Then step
# functions call schema_bind_state/schema_bind_input directly.
#
# HARD RULE: self_hosting/lib/test.patlang's run_feature never resets bound
# vars between scenarios (only t_skipping/t_pending_tags are per-scenario).
# Every scenario's Given/Then must explicitly rebind every declared state
# variable and input itself, every time -- including restating an unchanged
# variable in Then. Do not rely on a value surviving from a previous
# scenario, and do not assume an unmentioned variable defaults to its
# "before" value in Then -- an omitted rebind reads back as the not-yet-
# bound sentinel (see schema_state_values below), which will correctly
# fail the check rather than silently pass, but the failure will look like
# a real violation unless this rule is followed.
# =============================================================================
# =============================================================================
# pmap: a minimal Map<K,V> as an association list of [key, value] pairs with
# linear-scan lookup -- the exact same idiom already proven in
# self_hosting/lib/depgraph.patlang's depgraph_map_get/depgraph_map_append
# (a file->list-of-files map), generalised here to an arbitrary value type
# and given a full put/get/has/remove/keys surface. PatLang has no native
# Map/Dict type usable from self-hosted code at this scale.
#
# Functional / return-new-collection style throughout, for the same reason
# as pset.patlang: schema_bdd needs an independent before/after snapshot of
# state, which an in-place-mutating map wouldn't give for free.
#
# pmap_get returns [] (an empty list) for a missing key -- the same
# not-found sentinel depgraph_map_get already uses, not a distinct "Unit"
# literal (PatLang's dialect has no source-level Unit/nil literal to write
# directly). This is INHERENTLY AMBIGUOUS if a real stored value could
# itself be an empty list: pmap_has is the only call that actually
# distinguishes "absent" from "present but happens to look like the
# sentinel" -- never infer presence from pmap_get's return value alone.
# =============================================================================
make a function called pmap_new returns m
return []
end
make a function called pmap_has takes m, key returns found
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return true
end
let i = i + 1
end
return false
end
# See the module-level warning above: [] means "not found" here, which is
# indistinguishable from a genuinely stored empty-list value. Call
# pmap_has first whenever that distinction matters.
make a function called pmap_get takes m, key returns value
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return m[i][1]
end
let i = i + 1
end
return []
end
make a function called pmap_put takes m, key, value returns m2
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return list_set(m, i, [key, value])
end
let i = i + 1
end
return list_push(m, [key, value])
end
make a function called pmap_remove takes m, key returns m2
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] != key then
let out = list_push(out, m[i])
end
let i = i + 1
end
return out
end
make a function called pmap_keys takes m returns keys
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
let out = list_push(out, m[i][0])
let i = i + 1
end
return out
end
make a function called pmap_size takes m returns n
return to_num(list_len(m))
end
# Single, process-wide registry (like Step's own new("Step", text)
# convention -- schemas and operations are inherently global, not
# per-instance, so one lazily-created Dict is simpler than
# primitive_registry.patlang's counter-named multi-registry support, which
# this doesn't need).
make a function called schema_registry returns registry
let existing = get("__vars", "schema_bdd_registry_obj")
if existing then
return existing
end
let registry = new("Dict", "schema_bdd_registry")
set_var("schema_bdd_registry_obj", registry)
return registry
end
# state_var_names: list of strings naming the schema's declared state.
# invariant_fn: name of a function taking ONE argument (the state values,
# in the same order as state_var_names) and returning bool.
make a function called schema_define takes name, state_var_names, invariant_fn returns done
send(schema_registry(), "set", name + "__schema", [state_var_names, invariant_fn])
return true
end
make a function called schema_lookup takes name returns entry
return get(schema_registry(), name + "__schema")
end
# input_names: list of strings naming the operation's declared inputs.
# require_fn: name of a function taking (state_values, input_values),
# returning bool -- the operation's precondition.
# ensure_fn: name of a function taking (before_values, after_values,
# input_values), returning bool -- the operation's postcondition.
make a function called schema_operation takes schema_name, op_name, input_names, require_fn, ensure_fn returns done
send(schema_registry(), "set", schema_name + "::" + op_name + "__op", [schema_name, input_names, require_fn, ensure_fn])
return true
end
make a function called schema_lookup_operation takes schema_name, op_name returns entry
return get(schema_registry(), schema_name + "::" + op_name + "__op")
end
# ---- binding: a scenario's own step functions call these ----
make a function called schema_bind_state takes schema, var_name, phase, value returns done
set_var(schema + "__" + var_name + "__" + phase, value)
return true
end
make a function called schema_bind_input takes schema, op, param_name, value returns done
set_var(schema + "__" + op + "__in__" + param_name, value)
return true
end
# Not-yet-bound sentinel: get("__vars", ...) on an unset key -- same
# absence convention already relied on throughout this codebase (e.g.
# gherkin_contracts.patlang's `already_bound` check). Never distinguishes
# "never bound" from "bound to this same falsy value"; low risk for the
# LibraryLoans-shaped schemas this was designed against (string/list-typed
# state), a documented hard limit for anything boolean- or zero-valued.
make a function called schema_state_values takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__before"))
let i = i + 1
end
return out
end
make a function called schema_state_values_after takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__after"))
let i = i + 1
end
return out
end
make a function called schema_input_values takes schema, op, input_names returns values
let out = []
let i = 0
let n = to_num(list_len(input_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + op + "__in__" + input_names[i]))
let i = i + 1
end
return out
end
# ---- the witness check ----
#
# Four-stage check, each stage its own diagnosis tag, checked in the order
# a Z schema's own reasoning goes: is the state even valid to start from,
# does this operation's precondition actually hold for it, does the
# claimed resulting state stay valid, and does the operation's own
# postcondition connect before to after correctly. Returns [tag, payload],
# the same 2-element diagnosis shape as synth5_induce
# (self_hosting/lib/synthesis_lgg.patlang) -- deliberately, so
# schema_format_diagnosis below can mirror synth5_format_diagnosis's own
# real structure rather than inventing a new diagnosis style.
# The value-based core: takes state_before/inputs/state_after directly
# rather than reading them out of the __vars binding store. schema_check
# below is the BDD-scenario-shaped wrapper around this; this function is
# what any OTHER caller (a synthesis harness, a future GOAP or ILP hook --
# see the implementation plan's Phase 3) should call directly instead of
# faking a scenario's Given/When/Then bindings just to reach schema_check.
make a function called schema_check_values takes schema_name, op_name, state_before, inputs, state_after returns diagnosis
let schema_entry = schema_lookup(schema_name)
let invariant_fn = schema_entry[1]
let op_entry = schema_lookup_operation(schema_name, op_name)
let require_fn = op_entry[2]
let ensure_fn = op_entry[3]
if apply(invariant_fn, state_before) == false then
return ["invariant_violated_before", [schema_name, state_before]]
end
if apply(require_fn, state_before, inputs) == false then
return ["precondition_violated", [op_name, state_before, inputs]]
end
if apply(invariant_fn, state_after) == false then
return ["invariant_violated_after", [schema_name, state_after]]
end
if apply(ensure_fn, state_before, state_after, inputs) == false then
return ["postcondition_violated", [op_name, state_before, state_after, inputs]]
end
return ["ok", [op_name, state_before, state_after, inputs]]
end
make a function called schema_check takes schema_name, op_name returns diagnosis
let schema_entry = schema_lookup(schema_name)
let state_var_names = schema_entry[0]
let op_entry = schema_lookup_operation(schema_name, op_name)
let input_names = op_entry[1]
let state_before = schema_state_values(schema_name, state_var_names)
let inputs = schema_input_values(schema_name, op_name, input_names)
let state_after = schema_state_values_after(schema_name, state_var_names)
return schema_check_values(schema_name, op_name, state_before, inputs, state_after)
end
# Mirrors synth5_format_diagnosis's real structure (synthesis_lgg.patlang):
# switch on diagnosis[0], build a human-readable question per case. Returns
# a plain list of question strings (empty for "ok"), same shape as
# synth5_induce_and_ask's own return convention.
make a function called schema_format_diagnosis takes schema_name, op_name, diagnosis returns questions
let tag = diagnosis[0]
let payload = diagnosis[1]
if tag == "invariant_violated_before" then
return ["This scenario's Given already leaves " + schema_name + " in a state that violates its own invariant, before " + op_name + " is even checked. Is the Given wrong, or is the invariant too strict?"]
end
if tag == "precondition_violated" then
return ["Scenario claims " + op_name + " can run from this Given, but " + op_name + "'s own precondition returned false for these inputs. Is the Given wrong, or is " + op_name + "'s precondition too strict -- or is this scenario meant to test a rejection path, which needs its own operation schema rather than " + op_name + "'s?"]
end
if tag == "invariant_violated_after" then
return ["Applying " + op_name + " produces a state that violates " + schema_name + "'s invariant. Is the Then clause's claimed resulting state wrong, or does " + op_name + " need a stronger precondition to rule this case out?"]
end
if tag == "postcondition_violated" then
return ["The before/after state this scenario claims for " + op_name + " does not satisfy its own postcondition. Is the Then clause's claimed resulting state wrong, or is " + op_name + "'s postcondition wrong?"]
end
return []
end
# ---- synthesis integration (implementation plan Phase 3) ----
#
# Opt-in registration linking a SYNTHESIZED function's name to a schema
# operation it's meant to satisfy, plus a "harness" function name that
# knows how to actually exercise it: harness_fn takes (func_name) and
# returns [state_before, inputs, state_after] by calling the synthesized
# function itself (via apply) on some concrete test input and observing
# what it does. This is necessarily domain-specific -- there is no way to
# derive it generically -- so it must be supplied by whoever registers the
# hook, not synthesized here.
#
# Deliberately separate from schema_operation itself: a schema operation
# describes the CONTRACT; a synthesis hook additionally says "and here's
# how to actually run a candidate implementation against it," which only
# matters once there's a synthesized candidate to check, not for the
# scenario-witnessing use in schema_check above.
make a function called schema_register_synthesis_check takes func_name, schema_name, op_name, harness_fn returns done
send(schema_registry(), "set", func_name + "__synthesis_check", [schema_name, op_name, harness_fn])
return true
end
make a function called schema_lookup_synthesis_check takes func_name returns entry
return get(schema_registry(), func_name + "__synthesis_check")
end
# Runs a registered synthesis hook for func_name and returns its
# schema_check_values diagnosis directly. Callers with no hook registered
# for func_name should skip calling this entirely (schema_lookup_
# synthesis_check(func_name) is falsy) rather than call it and inspect the
# result -- there is no "no hook registered" diagnosis tag, since this
# function assumes a hook exists.
make a function called schema_run_synthesis_check takes func_name returns diagnosis
let hook = schema_lookup_synthesis_check(func_name)
let schema_name = hook[0]
let op_name = hook[1]
let harness_fn = hook[2]
let triple = apply(harness_fn, func_name)
return schema_check_values(schema_name, op_name, triple[0], triple[1], triple[2])
end
# mapper_fn: a caller-supplied function taking (path, resulting_facts) and
# returning [state_before, inputs, state_after] -- the same triple shape
# schema_check_values expects, built however the specific schema's domain
# needs from the plan's real path/resulting-facts.
#
# Returns ["no_plan_found", [goal_facts]] if plan_with_state finds no
# plan at all within its search cap -- a distinct diagnosis tag from
# schema_check_values' own four, since "no plan exists" is a planning
# failure, not a schema violation.
make a function called schema_check_goap_plan takes schema_name, op_name, goal_facts, mapper_fn returns diagnosis
let result = plan_with_state(goal_facts)
let path = result[0]
let resulting_facts = result[1]
if to_num(list_len(path)) == 0 then
return ["no_plan_found", [goal_facts]]
end
let triple = apply(mapper_fn, path, resulting_facts)
return schema_check_values(schema_name, op_name, triple[0], triple[1], triple[2])
end
# Membership check over a resulting_facts_list ([pred, args_list] pairs)
# -- the shape a schema's ensure/invariant predicate over GOAP state will
# usually need, since resulting_facts is itself the state, not a
# pset/pmap. Small enough to live here rather than grow pset.patlang's
# surface for a shape specific to this bridge.
make a function called schema_facts_contains takes facts, pred, args returns found
let i = 0
let n = to_num(list_len(facts))
while i < n do
if (facts[i][0] == pred) and (facts[i][1] == args) then
return true
end
let i = i + 1
end
return false
end
# ---- TransferPolicy schema: no general invariant (state = []); the real
# check is entirely in TransferBook's own postcondition, which reads the
# plan's ACTUAL resulting facts.
make a function called tp_inv_trivial takes state returns ok
return true
end
make a function called tp_require_true takes state, inputs returns ok
return true
end
# The real check: the book must end up at the destination branch, AND
# must not simultaneously still be recorded at the origin branch --
# a two-clause property no single action label could express on its own.
make a function called tp_ensure_arrived_and_gone takes before, after, inputs returns ok
let resulting_facts = after[0]
let book = inputs[0]
let origin = inputs[1]
let destination = inputs[2]
let arrived = schema_facts_contains(resulting_facts, "at", [book, destination])
let still_at_origin = schema_facts_contains(resulting_facts, "at", [book, origin])
return arrived and (still_at_origin == false)
end
make a function called tp_mapper takes path, resulting_facts returns triple
return [[resulting_facts], ["rare_atlas", "branch_a", "branch_b"], [resulting_facts]]
end
t_init()
schema_define("TransferPolicy", [], "tp_inv_trivial")
schema_operation("TransferPolicy", "TransferBook", ["book", "origin", "destination"], "tp_require_true", "tp_ensure_arrived_and_gone")
# ---- real GOAP domain: three actions moving a book through a depot.
action_add("pack_for_transit", [["at", ["Book", "branch_a"]]], [["in_transit", ["Book"]]], [["at", ["Book", "branch_a"]]], 1)
action_add("ship_to_depot", [["in_transit", ["Book"]]], [["at", ["Book", "depot"]]], [["in_transit", ["Book"]]], 2)
action_add("ship_to_branch", [["at", ["Book", "depot"]]], [["at", ["Book", "branch_b"]]], [["at", ["Book", "depot"]]], 2)
rule_add("at", ["rare_atlas", "branch_a"], [])
let transfer_goal = [["at", ["rare_atlas", "branch_b"]]]
let raw_result = plan_with_state(transfer_goal)
check("the planner finds the full three-hop route", to_num(list_len(raw_result[0])), 3)
check("step 1 packs the book for transit", raw_result[0][0], "pack_for_transit(Book=rare_atlas)")
check("step 2 ships it to the depot", raw_result[0][1], "ship_to_depot(Book=rare_atlas)")
check("step 3 ships it on to the destination branch", raw_result[0][2], "ship_to_branch(Book=rare_atlas)")
check("the real resulting state has the book at branch_b", schema_facts_contains(raw_result[1], "at", ["rare_atlas", "branch_b"]), true)
check("the real resulting state no longer has it at branch_a", schema_facts_contains(raw_result[1], "at", ["rare_atlas", "branch_a"]), false)
check("the real resulting state has no dangling in-transit record", schema_facts_contains(raw_result[1], "in_transit", ["rare_atlas"]), false)
let diagnosis = schema_check_goap_plan("TransferPolicy", "TransferBook", transfer_goal, "tp_mapper")
check("the full transfer plan satisfies the transfer policy", diagnosis[0], "ok")
# ---- an unreachable transfer: no route exists to a branch nothing ships to.
let impossible_goal = [["at", ["rare_atlas", "branch_c"]]]
let no_route_diagnosis = schema_check_goap_plan("TransferPolicy", "TransferBook", impossible_goal, "tp_mapper")
check("a destination with no shipping route is reported as no_plan_found", no_route_diagnosis[0], "no_plan_found")
t_report()
# Selftest for self_hosting/lib/schema_synthesis_bridge.patlang
# (implementation plan Phase 3, item 2): a REAL synth5_induce call derives
# grandparent(X) :- parent(X,Y), parent(Y,Z) from training facts and
# examples, exactly as in synthesis_lgg_selftest.patlang's own scenario A
# -- then the induced rule is checked against a GrandparentPolicy schema
# stating an independent constraint (a restricted-names list) the
# induction engine had no way to know about, since it only ever reasoned
# over parent/grandparent facts. One of the two candidates the induced
# rule proves true ("dave") is on the restricted list; the other
# ("alice") isn't -- proving the bridge catches a logically-correct
# induced rule producing a value an unrelated schema still forbids.
#
# Run from the repo root:
# rust-runtime/target/release/pat --ir-run self_hosting/schema_synthesis_bridge_selftest.patlang
# =============================================================================
# Test framework (Stage 1 dialect): unit assertions plus a Gherkin-style
# feature runner. Step definitions are registered in the object store keyed
# by their text; features are plain text dispatched line by line, so the
# same framework covers unit, integration, and behaviour tests.
# =============================================================================
make a function called t_init returns done
set_var("t_pass", 0)
set_var("t_fail", 0)
set_var("t_tagfilter", "")
set_var("t_pending_tags", "")
set_var("t_skipping", 0)
return true
end
make a function called contains_text takes hay, needle returns r
if needle.length > hay.length then
return false
end
let i = 0
while i <= hay.length - needle.length do
if substr(hay, i, needle.length) == needle then
return true
end
let i = i + 1
end
return false
end
make a function called check takes label, actual, expected returns ok
if actual == expected then
set_var("t_pass", get("__vars", "t_pass") + 1)
print(" ok: " + label)
return true
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: " + label + " (got " + actual + ", want " + expected + ")")
return false
end
end
make a function called t_report returns ok
let p = get("__vars", "t_pass")
let f = get("__vars", "t_fail")
print("tests: " + p + " passed, " + f + " failed")
if f == 0 then
print("ALL TESTS PASSED")
return true
else
print("TESTS FAILED")
return false
end
end
# ---- Gherkin runner ----
# Register a step: step("a fresh till", "st_fresh_till")
make a function called step takes text, fname returns done
new("Step", text)
send(text, "set", "fn", fname)
return true
end
make a function called starts_with takes s, prefix returns r
if s.length < prefix.length then
return false
end
return substr(s, 0, prefix.length) == prefix
end
make a function called trim_left takes s returns out
let i = 0
let scanning = true
while (i < s.length) and scanning do
let c = char_code(s, i)
if (c == 32) or (c == 9) then
let i = i + 1
else
let scanning = false
end
end
return substr(s, i, s.length - i)
end
# Strip a Gherkin keyword; returns the step text or "" if not a step line
make a function called step_text takes line returns out
if starts_with(line, "Given ") then
return substr(line, 6, line.length - 6)
end
if starts_with(line, "When ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "Then ") then
return substr(line, 5, line.length - 5)
end
if starts_with(line, "And ") then
return substr(line, 4, line.length - 4)
end
return ""
end
# Run only scenarios whose preceding @tag line contains `tag` ("" = all)
make a function called run_feature_tagged takes feature, tag returns ok
set_var("t_tagfilter", tag)
return run_feature(feature)
end
make a function called run_feature_file takes path returns ok
return run_feature(read_file(path))
end
make a function called run_feature takes feature returns ok
let h = str_intern(feature)
let n = sc_len(h)
let i = 0
let line = sb_new()
while i <= n do
let c = sc_code(h, i)
if (c == 10) or (c == -1) then
let raw = trim_left(sb_str(line))
let line = sb_new()
if starts_with(raw, "@") then
set_var("t_pending_tags", raw)
end
if starts_with(raw, "Feature:") then
print(raw)
end
if starts_with(raw, "Scenario:") then
let filter = get("__vars", "t_tagfilter")
let tags = get("__vars", "t_pending_tags")
set_var("t_pending_tags", "")
if filter then
if tags then
if contains_text(tags, filter) then
set_var("t_skipping", 0)
print(raw + " [" + tags + "]")
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 1)
print(raw + " [skipped: needs " + filter + "]")
end
else
set_var("t_skipping", 0)
print(raw)
end
else
let text = step_text(raw)
if (text != "") and (get("__vars", "t_skipping") != 1) then
if starts_with(text, "require ") then
handle_contract_step("require", substr(text, 8, text.length - 8))
else
if starts_with(text, "ensure ") then
handle_contract_step("ensure", substr(text, 7, text.length - 7))
else
let fname = get(text, "fn")
if fname then
apply(fname)
else
set_var("t_fail", get("__vars", "t_fail") + 1)
print(" FAIL: undefined step: " + text)
end
end
end
end
end
let i = i + 1
else
if c == 13 then
let i = i + 1
else
sb_push(line, sc_char(h, i))
let i = i + 1
end
end
end
return true
end
# =============================================================================
# pset: a minimal Set<T> built on a plain list with linear-scan membership,
# the same idiom already used by depgraph_list_contains and
# synth5_list_contains (self_hosting/lib/depgraph.patlang,
# self_hosting/lib/synthesis_lgg.patlang) rather than a new primitive --
# PatLang has no native Set type.
#
# Functional / return-new-collection style throughout (mirrors list_push's
# own semantics): every mutating-sounding operation returns a NEW list
# rather than mutating in place. This matters specifically for schema_bdd's
# before/after snapshotting, which needs an independent copy of state at
# two points in time -- an in-place structure would make "before" silently
# turn into "after" once the operation runs.
# =============================================================================
make a function called pset_new returns s
return []
end
make a function called pset_contains takes s, item returns found
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] == item then
return true
end
let i = i + 1
end
return false
end
make a function called pset_add takes s, item returns s2
if pset_contains(s, item) then
return s
end
return list_push(s, item)
end
make a function called pset_remove takes s, item returns s2
let out = []
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] != item then
let out = list_push(out, s[i])
end
let i = i + 1
end
return out
end
make a function called pset_size takes s returns n
return to_num(list_len(s))
end
# Identity -- exposed so callers iterating a set's members don't need to
# know it's "just a list" internally; the name documents intent at the
# call site instead.
make a function called pset_to_list takes s returns items
return s
end
# Set equality: same size, and every member of a is in b. Given pset_add's
# own dedup, equal size plus one-directional containment implies the
# reverse containment too (no way for b to hold something a doesn't
# without differing in size).
make a function called pset_equal takes a, b returns eq
if pset_size(a) != pset_size(b) then
return false
end
let i = 0
let n = to_num(list_len(a))
while i < n do
if pset_contains(b, a[i]) == false then
return false
end
let i = i + 1
end
return true
end
# =============================================================================
# Milestone 1 of the BDD-driven inductive synthesis system (see plan
# "wondering-about-extending-patlang-modular-shore"). Reads a toy feature
# text of the form
#
# Given the input is "1"
# Then the category is "low"
#
# scenario pairs, induces one PatLang `rule <category>(<input>).` fact per
# example, groups examples that share a category under the same predicate
# (LGG-by-output-label: the parsimony property under test is "branch count
# == real category count, not example count"), verifies every example
# against the induced rule set via the native solve() engine, then emits
# literal PatLang source text for the induced rules plus a small dispatcher
# function.
#
# No hand-rolled unification: induction/verification both go through the
# existing rule_add/solve host functions (self_hosting/examples/
# rule_syntax_demo.patlang), not a new resolution engine.
# =============================================================================
# ---- scenario extraction ----
# Return the text between the first pair of double-quotes on `line`, or ""
# if there isn't one.
make a function called synth_extract_quoted takes line returns out
let h = str_intern(line)
let n = sc_len(h)
let i = 0
let start = -1
let result = ""
while i < n do
if sc_code(h, i) == 34 then
if start == -1 then
let start = i + 1
else
let result = substr(line, start, i - start)
let i = n
end
end
let i = i + 1
end
return result
end
# Parse a toy feature text into a list of [input, category] pairs. Only
# understands the two step shapes this milestone's toy scenarios use
# (`Given the input is "X"` / `Then the category is "Y"`) -- deliberately
# narrow rather than a general Gherkin grammar; extend if a later example
# needs more.
make a function called synth_parse_examples takes feature returns examples
let h = str_intern(feature)
let n = sc_len(h)
let i = 0
let line = sb_new()
let pending_input = ""
let examples = []
while i <= n do
let c = sc_code(h, i)
if (c == 10) or (c == -1) then
let raw = sb_str(line)
let line = sb_new()
if contains_text(raw, "the input is") then
let pending_input = synth_extract_quoted(raw)
end
if contains_text(raw, "the category is") then
let category = synth_extract_quoted(raw)
if (pending_input != "") and (category != "") then
let examples = list_push(examples, [pending_input, category])
end
end
let i = i + 1
else
if c == 13 then
let i = i + 1
else
sb_push(line, sc_char(h, i))
let i = i + 1
end
end
end
return examples
end
# ---- induction ----
# The native A1 resolver treats any argument string matching `^[A-Z]` (an
# uppercase first letter) as a Prolog-style logic VARIABLE, not a ground
# constant (rust-runtime/src/ir/hosts.rs, `is_logic_var`) -- so an example
# value like "GET /users" would silently unify with anything if passed to
# rule_add/solve unprefixed. Ground every term with a lowercase marker
# before it touches the resolver so this convention can never misfire on
# example data, regardless of what the data itself looks like.
make a function called synth_ground takes term returns grounded
return "v:" + term
end
# True if `needle` already occurs in list `xs` (linear scan, string equality).
make a function called synth_list_contains takes xs, needle returns found
let i = 0
let n = to_num(list_len(xs))
let found = false
while i < n do
if xs[i] == needle then
let found = true
end
let i = i + 1
end
return found
end
# Register one ground fact rule (category(input) :- .) per example, and
# return the list of distinct categories seen -- the induced branch set.
# This is the LGG-by-output-label step: examples sharing a category become
# extra facts under one predicate instead of one predicate per example.
make a function called synth_induce takes examples returns categories
let i = 0
let n = to_num(list_len(examples))
let categories = []
while i < n do
let pair = examples[i]
let input = pair[0]
let category = pair[1]
rule_add(category, [synth_ground(input)], [])
if not synth_list_contains(categories, category) then
let categories = list_push(categories, category)
end
let i = i + 1
end
return categories
end
# ---- verification (CEGIS-lite: check every example against the induced
# rule set; return the count that fail) ----
make a function called synth_verify takes examples returns fail_count
let i = 0
let n = to_num(list_len(examples))
let fail_count = 0
while i < n do
let pair = examples[i]
let input = pair[0]
let category = pair[1]
let sols = solve(category, [synth_ground(input)])
if to_num(list_len(sols)) == 0 then
let fail_count = fail_count + 1
end
let i = i + 1
end
return fail_count
end
# ---- codegen: emit literal PatLang source for the induced rules plus a
# dispatcher function that tries each known category via solve() ----
make a function called synth_emit_rules takes examples returns src
let sb = sb_new()
let i = 0
let n = to_num(list_len(examples))
while i < n do
let pair = examples[i]
sb_push(sb, "rule " + pair[1] + "(\"" + synth_ground(pair[0]) + "\").\n")
let i = i + 1
end
return sb_str(sb)
end
make a function called synth_emit_dispatcher takes categories, fn_name returns src
let sb = sb_new()
sb_push(sb, "make a function called " + fn_name + " takes x returns category\n")
let i = 0
let n = to_num(list_len(categories))
while i < n do
let cat = categories[i]
sb_push(sb, " let sols_" + cat + " = solve(\"" + cat + "\", [\"v:\" + x])\n")
sb_push(sb, " if to_num(list_len(sols_" + cat + ")) > 0 then\n")
sb_push(sb, " return \"" + cat + "\"\n")
sb_push(sb, " end\n")
let i = i + 1
end
sb_push(sb, " return \"unknown\"\n")
sb_push(sb, "end\n")
return sb_str(sb)
end
# =============================================================================
# Milestone 5 of the BDD-driven inductive synthesis system (see plan
# "wondering-about-extending-patlang-modular-shore"). Milestones 2-4 are
# all TOP-DOWN template search: guess a clause shape (base+chain, unary
# conjunction, binary chain), enumerate every combination the template
# allows, and test each one blindly against the examples. None of them
# derive the shape FROM the examples -- they just brute-force a fixed
# hypothesis space.
#
# This module is BOTTOM-UP instead, closer to how real ILP systems
# (Aleph/Progol's "bottom clause" construction) actually work: for each
# positive example, walk the background-fact graph outward from that
# example's value to find a WITNESSED chain of relations connecting it to
# something -- a real piece of evidence, not a guess -- then generalize
# across examples by taking the longest common prefix of their witnessed
# chains (this prefix-taking IS anti-unification/LGG: the "terms" being
# generalized are predicate-sequences, and dropping the tail where they
# diverge is exactly "replace what differs with nothing further
# specified"). Only then is the generalized hypothesis verified against
# ALL examples (reusing milestone 2/4's subprocess-isolation machinery).
#
# The payoff for going bottom-up: when induction fails, WHY is directly
# visible from the evidence gathered along the way -- which positive
# examples had no witness at all (a genuine hole in the background
# knowledge, not just "search failed"), and which negative examples the
# generalized-too-far hypothesis wrongly covers. See synth5_induce's
# returned diagnosis.
#
# Background facts are taken here as real PatLang DATA (a list of
# [pred, arg1] / [pred, arg1, arg2] tuples), not hand-written rule-source
# text like milestones 2-4 required -- this module derives the rule
# source itself (synth5_facts_to_src), so there's exactly one
# representation of the background knowledge to keep consistent, not two.
#
# Milestone 6 (see plan "wondering-about-extending-patlang-modular-shore"):
# the witness search originally took the FIRST matching background fact
# at every hop (synth5_first_binary_hop/synth5_bfs_chain, kept below
# unchanged as a baseline) -- fine for a simple chain-shaped fact graph,
# wrong for a real one with actual branching (e.g. one parent with two
# children only ever witnessed whichever child's fact happened to be
# listed first). synth5_all_binary_hops/synth5_bfs_all_chains explore
# every witnessed branch instead, and synth5_lgg_from_positive now
# anti-unifies across each example's full chain SET rather than a single
# greedily-chosen chain -- with a new "no_common_structure" diagnosis for
# the case where every example has evidence, but that evidence never
# converges on one shared shape.
# =============================================================================
# =============================================================================
# Milestone 4 of the BDD-driven inductive synthesis system (see plan
# "wondering-about-extending-patlang-modular-shore"). Milestones 2 and 3
# both induced clauses over a fixed head shape `target(X) :- ...` where
# every body literal is applied to X itself (a unary conjunction) or to a
# single self-recursive call (the fixed base/chain template). Neither can
# express a genuine multi-hop RELATIONAL join through fresh existential
# variables -- the classic ILP benchmark:
# grandparent(X) :- parent(X, Y), parent(Y, Z).
# (Y and Z never appear in the head; they're existentially quantified,
# Prolog's ordinary "exists a Y and Z such that..." reading.)
#
# This module searches non-recursive chains of BINARY background
# predicates: `target(X) :- R1(X, Y1), R2(Y1, Y2), ..., Rk(Y(k-1), Yk).`
# over increasing depth k and candidate predicate assignment per hop
# (with repetition -- grandparent uses the same `parent` predicate twice),
# smallest depth first. Reuses milestone 2's subprocess-per-hypothesis
# query/test machinery unchanged (self_hosting/lib/
# synthesis_recursive.patlang's synth2_query/synth2_test_hypothesis).
# =============================================================================
# =============================================================================
# Milestone 2 of the BDD-driven inductive synthesis system (see plan
# "wondering-about-extending-patlang-modular-shore" and the memory note on
# why milestone 1 couldn't attempt this: milestone 1's LGG-by-output-label
# only groups flat ground facts under one predicate per category -- it has
# no way to induce a RECURSIVE clause body like
# `buildable(X) :- dep(X, Y), buildable(Y).` from buildable/unbuildable
# examples alone.
#
# This module does that, using a small fixed metarule library (Metagol-
# style meta-interpretive learning, not open-ended structural
# anti-unification -- that's still future work): given background facts,
# a target predicate name, and positive/negative example queries, it
# searches over combinations of
# base: target(X) :- Q(X). -- Q a background unary predicate
# chain: target(X) :- Q(X, Y), target(Y). -- Q a background binary predicate
# and accepts the first (smallest) combination whose induced rule set
# proves every positive example and fails every negative one.
#
# rule_add/RULES has no PatLang-exposed removal API, so a wrong hypothesis
# can't be un-registered from the current process without polluting later
# hypotheses. Each hypothesis is therefore tested in a clean interpreter
# (interp_run: an in-process world swap where the runtime has one, a fresh
# pat.exe otherwise) rather than via in-process rule_add/solve. A hypothesis
# whose test script fails to run counts as rejected, never as a pass.
# =============================================================================
# interp_run.patlang -- run PatLang source in a clean interpreter and get its
# output back as a value: [stdout, stderr, ok].
#
# Two backends behind one API:
# "world" in-process: compile with lexer/parser/lower, then the world_run
# host (every global store swapped for an empty one, print
# captured, caller's world restored afterwards). Works wherever
# host_caps() lists "world_swap" -- native pat.exe, the compiled
# runtime and the browser playground -- and needs no subprocess.
# "process" a fresh pat.exe --quiet --ir-run child via exec_capture_io.
# Needs "subprocess" in host_caps() and a pat.exe; honours a
# wall-clock timeout_ms and stdin.
#
# Options are a list of [key, value] pairs (all optional):
# backend "auto" (default) | "world" | "process"
# max_steps world only: loop-iteration/call budget; exhausted => ok=false
# timeout_ms process only: wall-clock limit; expiry => ok=false
# stdin process only: text fed to the child's stdin
# base_dir resolve `include` lines in the source relative to this
# directory before running (both backends; without it an
# include line is left for the child to fail on)
# pat_path path to pat.exe (else $PATLANG_PAT, else
# rust-runtime/target/release/pat.exe)
# vfs_in world only: [[path, content], ...] preloaded into the child
# vfs_out_prefix world only: only VFS paths under this prefix come back
#
# "auto" picks the world backend when it exists (a spawn costs orders of
# magnitude more than a swap), and the process backend when the caller asks
# for something only it can do (timeout_ms, stdin) or the world backend is
# missing (a native x64 program).
#
# Not supported by the world backend: `syntax { ... }` DSL blocks (expand
# them first), reading stdin.
# Self-hosted mirror of rust-runtime/src/preprocess.rs's expand_includes:
# expands `include "relative/path.patlang"` lines by splicing the
# referenced file's contents in place, resolving paths relative to the
# including file's own directory, recursively.
#
# Why this exists as PatLang, not just Rust: `expand_includes` was
# previously a NATIVE-ONLY preprocessing step (main.rs, run before the
# frontend ever sees the source) -- patc1.exe's own self-hosted lexer/
# parser never learned to do this, so any .patlang file using `include`
# could only be compiled via the native pat.exe frontend (`--ir-run`/
# `--patc`), never handed directly to patc1.exe, which is why every
# multi-file portfolio demo in build_portfolio.patlang manually
# concatenates dependency files (read_file(lexer) + chr(10) + ...) instead
# of using `include`. This closes that gap so `include` works identically
# everywhere -- interpreted, natively compiled, and self-hosted-compiled --
# matching this session's usual bar of "verified across all three paths."
#
# Kept as its own small library (not folded directly into patc1_main.patlang)
# so any self-hosted driver can `include "lib/includes.patlang"` and use it.
# str_trim(s) -> s with leading/trailing space/tab/\r/\n stripped.
# GitHub #22: \n was deliberately excluded here originally; audited every
# call site before changing this shared utility's semantics (per the
# issue's own explicit request not to "fix" it without checking callers
# first). Every current caller either trims an already-line-split string
# (no embedded \n to lose) or explicitly WANTS trailing newlines stripped
# (the two run_benchmarks.patlang/webcrawler.patlang callers comparing a
# captured-output blob across execution paths -- the original bug report
# that surfaced this: str_trim() alone didn't close a trailing-newline
# difference, needing a separate local helper to finish the job). No
# caller relies on \n being preserved through a trim call.
make a function called str_trim takes s returns trimmed
let n = s.length
let start = 0
while (start < n) and is_ws_char(char_code(s, start)) do
let start = start + 1
end
let end = n
while (end > start) and is_ws_char(char_code(s, end - 1)) do
let end = end - 1
end
return substr(s, start, end - start)
end
make a function called is_ws_char takes code returns is_ws
return (code == 32) or (code == 9) or (code == 13) or (code == 10)
end
# str_starts_with(s, prefix) -> bool
make a function called str_starts_with takes s, prefix returns matches
if prefix.length > s.length then
return false
end
return substr(s, 0, prefix.length) == prefix
end
# split_lines(s) -> list of lines, split on \n (a trailing \r on each line,
# from CRLF source files, is stripped too).
make a function called split_lines takes s returns lines
let out = []
let n = s.length
let start = 0
let i = 0
while i < n do
if char_code(s, i) == 10 then
let raw = substr(s, start, i - start)
let out = list_push(out, strip_trailing_cr(raw))
let start = i + 1
end
let i = i + 1
end
if start < n then
let out = list_push(out, strip_trailing_cr(substr(s, start, n - start)))
end
return out
end
make a function called strip_trailing_cr takes line returns stripped
let n = line.length
if (n > 0) and (char_code(line, n - 1) == 13) then
return substr(line, 0, n - 1)
end
return line
end
# path_dirname(path) -> everything before the last '/' or '\', or "." if
# the path has no directory component. Handles both separators since
# build_portfolio.patlang and friends run on Windows but write forward
# slashes in string literals.
make a function called path_dirname takes path returns dir
let n = path.length
let i = n - 1
let last_sep = -1
while i >= 0 do
let c = char_code(path, i)
if (c == 47) or (c == 92) then
let last_sep = i
let i = -1
else
let i = i - 1
end
end
if last_sep < 0 then
return "."
end
return substr(path, 0, last_sep)
end
# path_basename(path) -> everything after the last '/' or '\', or the
# whole path if it has no directory component -- the complement of
# path_dirname above (same separator-scanning loop, opposite half kept).
make a function called path_basename takes path returns base
let n = path.length
let i = n - 1
let last_sep = -1
while i >= 0 do
let c = char_code(path, i)
if (c == 47) or (c == 92) then
let last_sep = i
let i = -1
else
let i = i - 1
end
end
if last_sep < 0 then
return path
end
return substr(path, last_sep + 1, n - last_sep - 1)
end
# path_join(base, rel) -> base + "/" + rel, tolerating a trailing slash on
# base and an empty base (meaning "current directory"). If `rel` is itself
# absolute (leading '/'/'\', or a Windows drive letter like "C:"), it's
# returned unchanged, ignoring base -- matches Rust's PathBuf::join, which
# preprocess.rs's native expand_includes relies on for the same case.
make a function called path_join takes base, rel returns joined
if is_absolute_path(rel) then
return rel
end
if (base == "") or (base == ".") then
return rel
end
let n = base.length
if (n > 0) and ((char_code(base, n - 1) == 47) or (char_code(base, n - 1) == 92)) then
return base + rel
end
return base + "/" + rel
end
make a function called is_absolute_path takes p returns is_abs
if p.length == 0 then
return false
end
let c0 = char_code(p, 0)
if (c0 == 47) or (c0 == 92) then
return true
end
if (p.length >= 2) and (char_code(p, 1) == 58) then
return true
end
return false
end
# expand_includes(source, base_dir) -> source with every `include "path"`
# line recursively replaced by that file's own (recursively expanded)
# contents, paths resolved relative to base_dir (the including file's own
# directory) at each level, exactly matching preprocess.rs's semantics.
#
# The depth cap (16, matching preprocess.rs's MAX_DEPTH) is inlined as a
# literal below rather than a top-level `let` constant referenced from
# inside expand_includes_at_depth -- patc1.exe was found, while building
# this, to NOT make top-level `let` constants visible inside function
# bodies at all (confirmed via a minimal repro: the value silently reads
# as empty/unset, not an error) even though both --ir-run and native
# --patc handle this correctly. That's a real, previously-unknown
# self-hosted-compiler bug, logged separately in the backlog for its own
# dedicated fix -- this file just avoids relying on the broken behavior.
make a function called expand_includes takes source, base_dir returns expanded
return expand_includes_at_depth(source, base_dir, 0)
end
make a function called expand_includes_at_depth takes source, base_dir, depth returns expanded
if depth > 16 then
print("include: nesting deeper than 16 levels (cycle?)")
return source
end
let lines = split_lines(source)
let out = sb_new()
let i = 0
let n = to_num(list_len(lines))
while i < n do
let line = lines[i]
let t = str_trim(line)
if str_starts_with(t, "include ") and (str_starts_with(t, "#") == false) then
let rel = str_trim(substr(t, 8, t.length - 8))
let rel = strip_quotes(rel)
let path = path_join(base_dir, rel)
let inner = read_file(path)
let inner_base = path_dirname(path)
sb_push(out, expand_includes_at_depth(inner, inner_base, depth + 1))
sb_push(out, chr(10))
else
sb_push(out, line)
sb_push(out, chr(10))
end
let i = i + 1
end
return sb_str(out)
end
# strip_quotes("\"path\"") -> "path" -- include lines are always written
# with double-quoted paths, same as the native preprocessor expects.
make a function called strip_quotes takes s returns unquoted
let n = s.length
if (n >= 2) and (char_code(s, 0) == 34) and (char_code(s, n - 1) == 34) then
return substr(s, 1, n - 2)
end
return s
end
# =============================================================================
# Stage 1 self-hosted lexer library (Stage 0 compilable subset).
# Tokens are lists: [type, text, line] with types NUM, IDENT, STR, OP, NL, UNK, EOF.
# =============================================================================
make a function called is_digit_code takes c returns r
return (c >= 48) and (c <= 57)
end
make a function called is_alpha_code takes c returns r
if (c >= 65) and (c <= 90) then
return true
else
if (c >= 97) and (c <= 122) then
return true
else
return c == 95
end
end
end
make a function called is_op_code takes c returns r
if (c == 43) or (c == 45) or (c == 42) or (c == 47) or (c == 37) then
return true
else
if (c == 61) or (c == 60) or (c == 62) or (c == 33) then
return true
else
if (c == 40) or (c == 41) or (c == 91) or (c == 93) or (c == 44) or (c == 46) or (c == 124) or (c == 123) or (c == 125) or (c == 58) or (c == 59) then
return true
else
return false
end
end
end
end
make a function called tokenize takes src returns tokens
let tokens = vec_new()
let h = str_intern(src)
let n = sc_len(h)
let i = 0
let line = 1
while i < n do
let c = sc_code(h, i)
if c == 10 then
vec_push(tokens, ["NL", "", line])
let line = line + 1
let i = i + 1
else
if (c == 32) or (c == 9) or (c == 13) then
let i = i + 1
else
if c == 35 then
while (i < n) and (sc_code(h, i) != 10) do
let i = i + 1
end
else
if is_digit_code(c) then
# sb_new/sb_push/sb_str, NOT `txt = txt + ...` -- string
# concatenation in a loop copies the whole accumulated string
# on every append (PatLang strings are immutable), turning an
# O(n) scan into O(n^2). Numbers/identifiers are usually
# short so this rarely mattered in practice, but STRING
# literals (the branch below) can be arbitrarily long --
# self_hosting/lib/runtime_rs.patlang embeds large chunks of
# literal Rust source as single PatLang string literals,
# tens of thousands of characters each -- so this loop is a
# real, not just theoretical, O(n^2) risk on every compile
# that includes that file. Found via a new compiler warning
# (ir/lowering.rs) added specifically to catch this shape,
# while investigating an unrelated 30+-minute self-compile
# regression in a different file (self_hosting/lib/
# syntax_dsl.patlang) that turned out to have the identical
# anti-pattern.
let txt_b = sb_new()
let dots = 0
let scanning = true
while (i < n) and scanning do
let d = sc_code(h, i)
if is_digit_code(d) then
sb_push(txt_b, sc_char(h, i))
let i = i + 1
else
if (d == 46) and (dots == 0) then
let dots = 1
sb_push(txt_b, sc_char(h, i))
let i = i + 1
else
let scanning = false
end
end
end
vec_push(tokens, ["NUM", sb_str(txt_b), line])
else
if is_alpha_code(c) then
let txt_b = sb_new()
let scanning = true
while (i < n) and scanning do
let d = sc_code(h, i)
if is_alpha_code(d) or is_digit_code(d) then
sb_push(txt_b, sc_char(h, i))
let i = i + 1
else
let scanning = false
end
end
vec_push(tokens, ["IDENT", sb_str(txt_b), line])
else
if c == 34 then
let i = i + 1
let txt_b = sb_new()
let scanning = true
while (i < n) and scanning do
let d = sc_code(h, i)
if d == 34 then
let scanning = false
let i = i + 1
else
if d == 92 then
# escape sequences: n t r quote backslash (unknown kept raw)
let e = sc_code(h, i + 1)
if e == 110 then
sb_push(txt_b, chr(10))
let i = i + 2
else
if e == 116 then
sb_push(txt_b, chr(9))
let i = i + 2
else
if e == 114 then
sb_push(txt_b, chr(13))
let i = i + 2
else
if e == 34 then
sb_push(txt_b, chr(34))
let i = i + 2
else
if e == 92 then
sb_push(txt_b, chr(92))
let i = i + 2
else
sb_push(txt_b, sc_char(h, i))
let i = i + 1
end
end
end
end
end
else
sb_push(txt_b, sc_char(h, i))
let i = i + 1
end
end
end
vec_push(tokens, ["STR", sb_str(txt_b), line])
else
if is_op_code(c) then
let txt = sc_char(h, i)
let first = c
let i = i + 1
if i < n then
let d = sc_code(h, i)
if (d == 61) and ((first == 61) or (first == 33) or (first == 60) or (first == 62)) then
let txt = txt + sc_char(h, i)
let i = i + 1
else
# ':-' -- rule turnstile
if (d == 45) and (first == 58) then
let txt = txt + sc_char(h, i)
let i = i + 1
end
end
end
vec_push(tokens, ["OP", txt, line])
else
vec_push(tokens, ["UNK", sc_char(h, i), line])
let i = i + 1
end
end
end
end
end
end
end
end
vec_push(tokens, ["EOF", "", line])
return tokens
end
make a function called print_tokens takes tokens returns done
let n = vec_len(tokens)
let i = 0
while i < n do
let t = vec_get(tokens, i)
print(t[0] + " '" + t[1] + "' @" + t[2])
let i = i + 1
end
return true
end
# =============================================================================
# Stage 1 self-hosted parser library (Stage 0 compilable subset).
# Consumes tokens from lib/lexer.patlang, produces list-shaped AST nodes.
#
# Statements:
# ["Let", name, expr] let NAME = expr
# ["Expr", expr] call statements, e.g. print(x), emit(e, p)
# ["If", cond, [then], [else]] if expr then ... [else ...] end
# ["While", cond, [body]] while expr do ... end
# ["Func", name, [params], [b]] make a function called N takes a, b returns r ... end
# ["Return", expr] return expr
# ["When", event, [body], line] when EVENT do ... end (event handler)
# ["Err", message, line] parse error placeholder (1-indexed line)
#
# Expressions:
# ["Num", text] ["Str", text] ["Bool", "true"/"false"] ["Var", name]
# ["Bin", op, lhs, rhs] op: + - * / % == != < <= > >= and or
# ["Un", op, expr] op: not -
# ["Call", name, [args]]
# ["List", [items]]
# ["Index", obj, idx]
# ["Member", obj, prop]
#
# All parse functions return [node, next_pos] pairs; parse_args and
# parse_stmts_until return [list, next_pos].
# =============================================================================
make a function called tok_is takes t, ty, tx returns r
return (t[0] == ty) and (t[1] == tx)
end
make a function called skip_nl takes toks, pos returns p
# Skips newlines AND a bare ';' -- native parser.rs treats Semicolon as
# a fully optional statement separator right alongside Newline (never
# required; see its "while matches!(self.curr, Token::Semicolon |
# Token::Newline | ...)" loop), so this self-hosted mirror needs to
# tolerate the same already-existing, already-optional token, not
# introduce any new requirement into the grammar.
let p = pos
let looping = true
while looping do
let t = vec_get(toks, p)
if t[0] == "NL" then
let p = p + 1
else
if tok_is(t, "OP", ";") then
let p = p + 1
else
let looping = false
end
end
end
return p
end
# ---- expressions ----
make a function called parse_args takes toks, pos returns r
# pos points just after '('; returns [args, pos_after_rparen]
# tolerates newlines around '(', ',', and ')' so multi-line calls parse
let args = []
let p = skip_nl(toks, pos)
let t = vec_get(toks, p)
if tok_is(t, "OP", ")") then
return [args, p + 1]
else
let looping = true
while looping do
let e = parse_expr(toks, p)
let args = list_push(args, e[0])
let p = skip_nl(toks, e[1])
let t2 = vec_get(toks, p)
if tok_is(t2, "OP", ",") then
let p = skip_nl(toks, p + 1)
else
let looping = false
end
end
let t3 = vec_get(toks, p)
if tok_is(t3, "OP", ")") then
return [args, p + 1]
else
return [[["Err", "expected ) in argument list", vec_get(toks, p)[2]]], p]
end
end
end
make a function called parse_primary takes toks, pos returns r
let t = vec_get(toks, pos)
let ty = t[0]
let tx = t[1]
if ty == "NUM" then
return [["Num", tx], pos + 1]
else
if ty == "STR" then
return [["Str", tx], pos + 1]
else
if ty == "IDENT" then
if (tx == "true") or (tx == "false") then
return [["Bool", tx], pos + 1]
else
if (tx == "pursue") and (tok_is(vec_get(toks, pos + 1), "OP", "(") == false) then
# `pursue GOAL` -- GOAL is a bare name, taken as a string
# literal argument to the `pursue` host fn, mirroring native
# parser.rs's parse_primary special case exactly (the
# parenthesized form `pursue(x)` falls through to ordinary
# call parsing below instead).
let nameTok = vec_get(toks, pos + 1)
if nameTok[0] == "IDENT" then
return [["Call", "pursue", [["Str", nameTok[1]]]], pos + 2]
else
return [["Err", "expected goal name after 'pursue'", nameTok[2]], pos + 1]
end
else
if (tx == "activate") and (tok_is(vec_get(toks, pos + 1), "OP", "(") == false) then
# `activate PLAN` -- PLAN is a general expression (typically a
# variable holding a previous `pursue` result), unlike
# `pursue`'s bare goal-name.
let planR = parse_expr(toks, pos + 1)
return [["Call", "activate", [planR[0]]], planR[1]]
else
if (tx == "budgeted") and tok_is(vec_get(toks, pos + 1), "OP", "(") then
let msR = parse_expr(toks, pos + 2)
let msNode = msR[0]
let p = msR[1]
let existingNode = ["Bool", "false"]
let p2 = p
if tok_is(vec_get(toks, p), "OP", ",") then
let exR = parse_expr(toks, p + 1)
let existingNode = exR[0]
let p2 = exR[1]
end
if tok_is(vec_get(toks, p2), "OP", ")") then
let p3 = p2 + 1
let opener = vec_get(toks, p3)
if tok_is(opener, "OP", "{") then
let bodyR = parse_stmts_until(toks, p3 + 1)
let p4 = bodyR[1]
if tok_is(vec_get(toks, p4), "OP", "}") then
return [["Budgeted", msNode, existingNode, bodyR[0]], p4 + 1]
else
return [["Err", "expected } after budgeted body", vec_get(toks, p4)[2]], p4]
end
else
if (opener[0] == "IDENT") and (opener[1] == "do") then
let bodyR = parse_stmts_until(toks, p3 + 1)
let p4 = bodyR[1]
if tok_is(vec_get(toks, p4), "IDENT", "end") then
return [["Budgeted", msNode, existingNode, bodyR[0]], p4 + 1]
else
return [["Err", "expected end after budgeted body", vec_get(toks, p4)[2]], p4]
end
else
return [["Err", "expected { or do after budgeted(...)", vec_get(toks, p3)[2]], p3]
end
end
else
return [["Err", "expected ) after budgeted arguments", vec_get(toks, p2)[2]], p2]
end
else
let nx = vec_get(toks, pos + 1)
if tok_is(nx, "OP", "(") then
let a = parse_args(toks, pos + 2)
return [["Call", tx, a[0]], a[1]]
else
return [["Var", tx], pos + 1]
end
end
end
end
end
else
if tok_is(t, "OP", "(") then
let inner = parse_expr(toks, pos + 1)
let p = inner[1]
let t2 = vec_get(toks, p)
if tok_is(t2, "OP", ")") then
return [inner[0], p + 1]
else
return [["Err", "expected )", vec_get(toks, p)[2]], p]
end
else
if tok_is(t, "OP", "[") then
# tolerates newlines around '[', ',', and ']' for multi-line lists
let items = []
let p = skip_nl(toks, pos + 1)
let t2 = vec_get(toks, p)
if tok_is(t2, "OP", "]") then
return [["List", items], p + 1]
else
let looping = true
while looping do
let e = parse_expr(toks, p)
let items = list_push(items, e[0])
let p = skip_nl(toks, e[1])
let t3 = vec_get(toks, p)
if tok_is(t3, "OP", ",") then
let p = skip_nl(toks, p + 1)
else
let looping = false
end
end
let t4 = vec_get(toks, p)
if tok_is(t4, "OP", "]") then
return [["List", items], p + 1]
else
return [["Err", "expected ] in list", vec_get(toks, p)[2]], p]
end
end
else
if tok_is(t, "OP", "|") then
# Closure literal: |params| do body end, or |params| { body }
let p = pos + 1
let params = []
let looping = true
while looping do
let pt = vec_get(toks, p)
if tok_is(pt, "OP", "|") then
let p = p + 1
let looping = false
else
if pt[0] == "IDENT" then
let params = list_push(params, pt[1])
let p = p + 1
let nt = vec_get(toks, p)
if tok_is(nt, "OP", ",") then
let p = p + 1
end
else
let looping = false
end
end
end
let p = skip_nl(toks, p)
let dt = vec_get(toks, p)
if (dt[0] == "IDENT") and (dt[1] == "do") then
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
let et = vec_get(toks, p2)
if (et[0] == "IDENT") and (et[1] == "end") then
return [["Closure", params, body[0]], p2 + 1]
else
return [["Err", "expected end after closure body", vec_get(toks, p2)[2]], p2]
end
else
if tok_is(dt, "OP", "{") then
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
let et = vec_get(toks, p2)
if tok_is(et, "OP", "}") then
return [["Closure", params, body[0]], p2 + 1]
else
return [["Err", "expected } after closure body", vec_get(toks, p2)[2]], p2]
end
else
return [["Err", "expected do or { after closure params", vec_get(toks, p)[2]], p]
end
end
else
if tok_is(t, "OP", "{") then
# A bare `{ ... }` in expression position is a real,
# deliberate zero-param function literal -- the same
# value-position meaning `|params| { }` has with an empty
# param list, e.g. `let f = { return 42 }` then `f()`.
# Mirrors native parser.rs's parse_primary (Token::BlockStart)
# exactly. NOT legal as its own statement (see parse_stmt's
# dedicated rejection) -- only as a value.
let body = parse_stmts_until(toks, pos + 1)
let p2 = body[1]
let et = vec_get(toks, p2)
if tok_is(et, "OP", "}") then
return [["Closure", [], body[0]], p2 + 1]
else
return [["Err", "expected } after block body", vec_get(toks, p2)[2]], p2]
end
else
return [["Err", "unexpected " + ty + " '" + tx + "'", vec_get(toks, pos)[2]], pos + 1]
end
end
end
end
end
end
end
end
make a function called parse_postfix takes toks, pos returns r
let r1 = parse_primary(toks, pos)
let node = r1[0]
let p = r1[1]
let looping = true
while looping do
let t = vec_get(toks, p)
if tok_is(t, "OP", "[") then
let idx = parse_expr(toks, p + 1)
let p2 = idx[1]
let t2 = vec_get(toks, p2)
if tok_is(t2, "OP", "]") then
let node = ["Index", node, idx[0]]
let p = p2 + 1
else
let node = ["Err", "expected ] after index", vec_get(toks, p2)[2]]
let looping = false
end
else
if tok_is(t, "OP", ".") then
let nameTok = vec_get(toks, p + 1)
if nameTok[0] == "IDENT" then
let node = ["Member", node, nameTok[1]]
let p = p + 2
else
let node = ["Err", "expected name after .", nameTok[2]]
let looping = false
end
else
# obj.method(args) -- a call whose callee is a Member expression,
# e.g. AST.RegisterRoute(...) (the RouterDSL example's own return
# template). The existing "Call" AST shape (used for bare
# NAME(args)) always treats its callee as a plain STRING, resolved
# by name against known functions/locals in lower.patlang -- there
# was no way to compose a Call around an arbitrary expression the
# way native parser.rs's Expr::Call{function: Box<Expr>, args}
# does. A dedicated MethodCall node (object expr + method name +
# args) sidesteps that rather than overloading "Call"'s callee
# slot with two different shapes; lower.patlang routes it to
# send(object, "method", ...args), the same host call obj.prop =
# value already uses on the write side.
if (node[0] == "Member") and tok_is(t, "OP", "(") then
let a = parse_args(toks, p + 1)
let node = ["MethodCall", node[1], node[2], a[0]]
let p = a[1]
else
let looping = false
end
end
end
end
return [node, p]
end
make a function called parse_unary takes toks, pos returns r
let t = vec_get(toks, pos)
if tok_is(t, "IDENT", "not") then
let e = parse_unary(toks, pos + 1)
return [["Un", "not", e[0]], e[1]]
else
if tok_is(t, "IDENT", "bnot") then
let e = parse_unary(toks, pos + 1)
return [["Un", "bnot", e[0]], e[1]]
else
if tok_is(t, "OP", "-") then
let e = parse_unary(toks, pos + 1)
return [["Un", "-", e[0]], e[1]]
else
return parse_postfix(toks, pos)
end
end
end
end
# GitHub #28: the binary-operator precedence chain used to be seven
# hand-cascaded functions (parse_mul -> parse_shift -> parse_add ->
# parse_bitwise -> parse_cmp -> parse_and -> parse_expr's own `or`
# tier), each re-implementing its OWN copy of "peek for a continuation
# operator, optionally past a newline, recurse." Every one of those
# seven copies was a separate place to forget the fix, which is
# exactly what happened (`let x =\n "a"` and several tiers were still
# missing it even after `+`/`-`/`*`/`/`/`%` had already been patched).
# Replaced with ONE table-driven precedence-climbing (Pratt) loop,
# `parse_binary`, so there is only one place left where "does this
# expression continue" is decided at all -- matching the native Rust
# frontend's own architecture (a single unified `parse_expression`,
# not a cascade), not just its behavior.
#
# binop_info(t) -> ["", 0, false] if t isn't a binary operator at all,
# else [canonical_op_name, precedence (1=loosest .. 7=tightest),
# continues_across_a_newline]. Precedence numbers match the OLD
# cascade's own nesting exactly (parse_expr=or=1 was outermost/
# loosest, parse_mul=7 was innermost/tightest) so existing precedence
# stays identical; only the MECHANISM collapsed to one function.
# `continues_across_a_newline` mirrors the native frontend's own
# whitelist (parser.rs's parse_expression) exactly: every arithmetic/
# comparison/bitwise/shift operator continues across a newline: `and`/
# `or` deliberately do NOT (matching that this is a real, general
# design DECISION already made by the reference implementation, not
# an oversight to "complete" -- `let ok = a\nand b` stays two separate
# things, same as it always has been in both frontends).
make a function called binop_info takes t returns info
if t[0] == "OP" then
if t[1] == "*" then
return ["*", 7, true]
end
if t[1] == "/" then
return ["/", 7, true]
end
if t[1] == "%" then
return ["%", 7, true]
end
if t[1] == "+" then
return ["+", 5, true]
end
if t[1] == "-" then
return ["-", 5, true]
end
if t[1] == "==" then
return ["==", 3, true]
end
if t[1] == "!=" then
return ["!=", 3, true]
end
if t[1] == "<" then
return ["<", 3, true]
end
if t[1] == "<=" then
return ["<=", 3, true]
end
if t[1] == ">" then
return [">", 3, true]
end
if t[1] == ">=" then
return [">=", 3, true]
end
return ["", 0, false]
else
if t[0] == "IDENT" then
if t[1] == "shl" then
return ["shl", 6, true]
end
if t[1] == "shr" then
return ["shr", 6, true]
end
if t[1] == "band" then
return ["band", 4, true]
end
if t[1] == "bxor" then
return ["bxor", 4, true]
end
if t[1] == "bor" then
return ["bor", 4, true]
end
if t[1] == "and" then
return ["and", 2, false]
end
if t[1] == "or" then
return ["or", 1, false]
end
return ["", 0, false]
else
return ["", 0, false]
end
end
end
make a function called parse_binary takes toks, pos, min_prec returns r
let r1 = parse_unary(toks, pos)
let node = r1[0]
let p = r1[1]
let looping = true
while looping do
# Same-line operator first (the common case, no newline involved
# at all); only if that's absent do we look PAST any newlines to
# see whether a continuation-eligible operator starts the next
# line -- this order is what keeps `let x = 5\n-3` two separate
# things: a leading `-` past a newline IS eligible in principle,
# but only actually taken if `binop_info`'s continues flag says so
# (true for `-`, matching the native frontend's own choice; this
# is an intentional, documented tradeoff shared with every C-like
# language that allows optional statement separators, not a bug).
let info = binop_info(vec_get(toks, p))
let oppos = p
if info[0] == "" then
let p2 = skip_nl(toks, p)
let info2 = binop_info(vec_get(toks, p2))
if info2[2] then
let info = info2
let oppos = p2
end
end
if (info[0] != "") and (info[1] >= min_prec) then
let rhs_start = skip_nl(toks, oppos + 1)
let r2 = parse_binary(toks, rhs_start, info[1] + 1)
let node = ["Bin", info[0], node, r2[0]]
let p = r2[1]
else
let looping = false
end
end
return [node, p]
end
make a function called parse_expr takes toks, pos returns r
# GitHub #28: an expression may legitimately START on the line AFTER
# whatever introduces it (`let x =\n "..."`, `return\n expr`, an
# `if`/`while` condition on its own line, etc.) -- skip any leading
# newlines here ONCE, at parse_expr's own single entry point, so
# every one of its call sites gets this for free, matching the
# native Rust frontend's own parse_expression (skips newlines before
# parsing its first primary, unconditionally, before the Pratt loop
# even starts).
return parse_binary(toks, skip_nl(toks, pos), 1)
end
# ---- statements ----
make a function called at_block_stop takes toks, pos returns r
let t = vec_get(toks, pos)
if t[0] == "EOF" then
return true
else
if (t[0] == "IDENT") and ((t[1] == "end") or (t[1] == "else") or (t[1] == "elif")) then
return true
else
if (t[0] == "OP") and (t[1] == "}") then
return true
else
return false
end
end
end
end
make a function called parse_stmts_until takes toks, pos returns r
# Collect statements until 'end' / 'else' / EOF (stopper not consumed).
let stmts = []
let p = skip_nl(toks, pos)
let looping = true
while looping do
if at_block_stop(toks, p) then
let looping = false
else
let s = parse_stmt(toks, p)
let stmts = list_push(stmts, s[0])
let p = skip_nl(toks, s[1])
end
end
return [stmts, p]
end
make a function called parse_if_then_tail takes toks, cond, pos, line returns r
# pos points just after 'then'. Handles [elif cond then ...]* [else ...] end.
# `line` is the line of the 'if'/'elif' keyword that introduced THIS
# clause (each elif in a chain desugars to its own nested If node with
# its own line, not the original 'if''s).
let thenPart = parse_stmts_until(toks, pos)
let p2 = thenPart[1]
let t2 = vec_get(toks, p2)
if tok_is(t2, "IDENT", "else") then
let elsePart = parse_stmts_until(toks, p2 + 1)
let p3 = elsePart[1]
let t3 = vec_get(toks, p3)
if tok_is(t3, "IDENT", "end") then
return [["If", cond, thenPart[0], elsePart[0], line], p3 + 1]
else
return [["Err", "expected end after else block", vec_get(toks, p3)[2]], p3]
end
else
if tok_is(t2, "IDENT", "elif") then
let elif_line = t2[2]
let c2 = parse_expr(toks, p2 + 1)
let cond2 = c2[0]
let p3 = c2[1]
let t3 = vec_get(toks, p3)
if tok_is(t3, "IDENT", "then") then
let nested = parse_if_then_tail(toks, cond2, p3 + 1, elif_line)
return [["If", cond, thenPart[0], [nested[0]], line], nested[1]]
else
return [["Err", "expected then after elif condition", vec_get(toks, p3)[2]], p3]
end
else
if tok_is(t2, "IDENT", "end") then
return [["If", cond, thenPart[0], [], line], p2 + 1]
else
return [["Err", "expected else, elif, or end after if block", vec_get(toks, p2)[2]], p2]
end
end
end
end
make a function called parse_if_brace_tail takes toks, cond, pos, line returns r
# pos points just after '{'. Handles [elif cond { ... }]* [else { ... }] '}'.
let thenPart = parse_stmts_until(toks, pos)
let p2 = thenPart[1]
let t2 = vec_get(toks, p2)
if tok_is(t2, "OP", "}") then
let p3 = p2 + 1
let t3 = vec_get(toks, p3)
if tok_is(t3, "IDENT", "else") then
let t4 = vec_get(toks, p3 + 1)
if tok_is(t4, "OP", "{") then
let elsePart = parse_stmts_until(toks, p3 + 2)
let p5 = elsePart[1]
let t5 = vec_get(toks, p5)
if tok_is(t5, "OP", "}") then
return [["If", cond, thenPart[0], elsePart[0], line], p5 + 1]
else
return [["Err", "expected } after else block", vec_get(toks, p5)[2]], p5]
end
else
return [["Err", "expected { after else", vec_get(toks, p3 + 1)[2]], p3 + 1]
end
else
if tok_is(t3, "IDENT", "elif") then
let elif_line = t3[2]
let c2 = parse_expr(toks, p3 + 1)
let cond2 = c2[0]
let p4 = c2[1]
let t4 = vec_get(toks, p4)
if tok_is(t4, "OP", "{") then
let nested = parse_if_brace_tail(toks, cond2, p4 + 1, elif_line)
return [["If", cond, thenPart[0], [nested[0]], line], nested[1]]
else
return [["Err", "expected { after elif condition", vec_get(toks, p4)[2]], p4]
end
else
return [["If", cond, thenPart[0], [], line], p3]
end
end
else
return [["Err", "expected } after if block", vec_get(toks, p2)[2]], p2]
end
end
make a function called parse_if takes toks, pos returns r
# pos points at 'if'. Supports both `if cond then ... end` and
# `if cond { ... }`, with `elif` accepted in either form.
# Line-tracking note: tokens already carry [type, text, line]
# (lexer.patlang) but If/While AST nodes previously dropped it. The
# 'if'/'elif'/'while' keyword's own line is threaded through as the
# LAST element of the node (append-only, so every existing consumer
# that reads node[1]/node[2]/node[3] by fixed index is unaffected) --
# added for self_hosting/lib/coverage.patlang's branch catalog and
# source-line instrumentation, see that file's header for why.
let if_line = vec_get(toks, pos)[2]
let c = parse_expr(toks, pos + 1)
let cond = c[0]
let p = c[1]
let t = vec_get(toks, p)
if tok_is(t, "IDENT", "then") then
return parse_if_then_tail(toks, cond, p + 1, if_line)
else
if tok_is(t, "OP", "{") then
return parse_if_brace_tail(toks, cond, p + 1, if_line)
else
return [["Err", "expected then or { after if condition", vec_get(toks, p)[2]], p]
end
end
end
make a function called parse_while takes toks, pos returns r
# pos points at 'while'. Supports both `while cond do ... end` and
# `while cond { ... }`. See parse_if's comment for why the keyword's
# line is appended as the node's last element.
let while_line = vec_get(toks, pos)[2]
let c = parse_expr(toks, pos + 1)
let cond = c[0]
let p = c[1]
let t = vec_get(toks, p)
if tok_is(t, "IDENT", "do") or tok_is(t, "IDENT", "begin") then
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
let t2 = vec_get(toks, p2)
if tok_is(t2, "IDENT", "end") then
return [["While", cond, body[0], while_line], p2 + 1]
else
return [["Err", "expected end after while body", vec_get(toks, p2)[2]], p2]
end
else
if tok_is(t, "OP", "{") then
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
let t2 = vec_get(toks, p2)
if tok_is(t2, "OP", "}") then
return [["While", cond, body[0], while_line], p2 + 1]
else
return [["Err", "expected } after while body", vec_get(toks, p2)[2]], p2]
end
else
return [["Err", "expected do or { after while condition", vec_get(toks, p)[2]], p]
end
end
end
# [message, line] -- message is "" if `nameTok` is fine to use as a
# function name, else the reserved-name error text. Shared by both
# function-definition surface syntaxes (`make a function called` and
# `fn`) since the native-codegen collision this guards against applies
# regardless of which one defined the function.
make a function called pf_check_main_reserved takes nameTok returns r
if nameTok[1] == "main" then
# "main" collides with the native codegen backend's own Rust `fn
# main()` entry point -- a user function literally named this
# compiled successfully via patc1.exe with no error, but the
# emitted call recursed into itself infinitely at runtime instead
# of dispatching to the real program entry, overflowing the stack.
#
# This is genuinely ONLY a native-compilation problem (--patc /
# patc1.exe) -- the tree-walking interpreter (--ir-run) has no Rust
# entry point to collide with at all, `main` there is just an
# ordinary function name. But the parser has no idea, at parse
# time, which backend the resulting AST will end up running on --
# a single parsed program can be hand to either. Rejecting `main`
# unconditionally here, for BOTH backends, is a deliberate, known
# over-restriction chosen for simplicity (one check, one place,
# catchable with a clear message at parse time) over a more
# precise but architecturally heavier fix (deferring the check to
# codegen time, backend-conditional). Worth relaxing later if it
# ever becomes a real nuisance for --ir-run-only code, but the
# honest long-term fix is making native codegen emit a Rust entry
# point name that genuinely cannot collide with any PatLang
# identifier, removing the need for this restriction altogether.
return ["'main' is a reserved function name", nameTok[2]]
end
return ["", 0]
end
make a function called parse_function_def takes toks, pos returns r
# pos points at 'make'; expect: make a function called NAME
# [takes a, b] [returns r] NL body end
let p = pos + 1
if tok_is(vec_get(toks, p), "IDENT", "a") then
let p = p + 1
end
if tok_is(vec_get(toks, p), "IDENT", "function") then
let p = p + 1
else
return [["Err", "expected 'function' after make", vec_get(toks, p)[2]], p]
end
if tok_is(vec_get(toks, p), "IDENT", "called") then
let p = p + 1
end
let nameTok = vec_get(toks, p)
if nameTok[0] == "IDENT" then
let name = nameTok[1]
let mainErr = pf_check_main_reserved(nameTok)
if mainErr[0] != "" then
return [["Err", mainErr[0], mainErr[1]], p]
end
let p = p + 1
let params = []
if tok_is(vec_get(toks, p), "IDENT", "takes") then
let p = p + 1
let looping = true
while looping do
let t = vec_get(toks, p)
if t[0] == "IDENT" then
if (t[1] == "returns") then
let looping = false
else
let params = list_push(params, t[1])
let p = p + 1
end
else
if tok_is(t, "OP", ",") then
let p = p + 1
else
let looping = false
end
end
end
end
# Named-return hint (GitHub issue #5): `returns r` gives r real
# fall-through semantics via lower.patlang's Func-lowering, matching
# Stage 0's rust-runtime/src/parser.rs:781-797 -- previously parsed
# and unconditionally discarded here (Stage 1 always required an
# explicit return). "" means no hint, matching the convention no
# other Func-node consumer treats an empty string specially.
let return_hint = ""
if tok_is(vec_get(toks, p), "IDENT", "returns") then
let hintTok = vec_get(toks, p + 1)
if hintTok[0] == "IDENT" then
let return_hint = hintTok[1]
end
let p = p + 2
end
# Body may be `... end` (word form) or `{ ... }` (brace form).
if tok_is(vec_get(toks, p), "OP", "{") then
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
if tok_is(vec_get(toks, p2), "OP", "}") then
return [["Func", name, params, body[0], return_hint], p2 + 1]
else
return [["Err", "expected } after function body", vec_get(toks, p2)[2]], p2]
end
else
let body = parse_stmts_until(toks, p)
let p2 = body[1]
if tok_is(vec_get(toks, p2), "IDENT", "end") then
return [["Func", name, params, body[0], return_hint], p2 + 1]
else
return [["Err", "expected end after function body", vec_get(toks, p2)[2]], p2]
end
end
else
return [["Err", "expected function name", vec_get(toks, p)[2]], p]
end
end
# `fn` form: fn NAME ( param, param, ... ) { body } -- shorter surface
# syntax for the same ["Func", name, params, body] node
# parse_function_def produces. No takes/returns keywords, and (unlike
# parse_function_def, which accepts either `{ }` or `... end`) the body
# is brace-only here -- matches native's parse_function
# (rust-runtime/src/parser.rs:706), which unconditionally expects '{'
# right after the parameter list, with no `end`-delimited alternative.
make a function called parse_fn_def takes toks, pos returns r
# pos points at 'fn'
let p = pos + 1
let nameTok = vec_get(toks, p)
if nameTok[0] != "IDENT" then
return [["Err", "expected function name", nameTok[2]], p]
end
let name = nameTok[1]
let mainErr = pf_check_main_reserved(nameTok)
if mainErr[0] != "" then
return [["Err", mainErr[0], mainErr[1]], p]
end
let p = p + 1
if tok_is(vec_get(toks, p), "OP", "(") == false then
return [["Err", "expected ( after function name", vec_get(toks, p)[2]], p]
end
let p = p + 1
let params = []
if tok_is(vec_get(toks, p), "OP", ")") == false then
let looping = true
while looping do
let t = vec_get(toks, p)
if t[0] == "IDENT" then
let params = list_push(params, t[1])
let p = p + 1
let t2 = vec_get(toks, p)
if tok_is(t2, "OP", ",") then
let p = p + 1
else
let looping = false
end
else
return [["Err", "expected parameter name", t[2]], p]
end
end
end
if tok_is(vec_get(toks, p), "OP", ")") == false then
return [["Err", "expected ) to close parameter list", vec_get(toks, p)[2]], p]
end
let p = p + 1
if tok_is(vec_get(toks, p), "OP", "{") == false then
return [["Err", "expected { to start function body", vec_get(toks, p)[2]], p]
end
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
if tok_is(vec_get(toks, p2), "OP", "}") then
return [["Func", name, params, body[0]], p2 + 1]
else
return [["Err", "expected } after function body", vec_get(toks, p2)[2]], p2]
end
end
# ---- match/case pattern matching (issue #44) ----
#
# Grammar:
# match EXPR do
# case PATTERN [when GUARD] then STMTS
# ...
# case _ then STMTS
# end
#
# Pattern AST shapes (never seen outside the parser/lowerer -- lower.patlang's
# compile_pattern is the only consumer):
# ["PWild"] _
# ["PBind", name] bare lowercase identifier
# ["PLit", litNode] Num/Str/Bool literal
# ["PCmp", op, exprNode] > < >= <= == != EXPR (a guard wearing pattern
# clothing -- see lower.patlang's compile_pattern)
# ["PGlob", globStr] *fred*, fred*, etc (unquoted, translated to a
# regex and matched via regex.patlang's glob_match)
# ["PList", [subpatterns]] [p1, p2, ...] -- recursive tagged-list pattern
#
# `match` is pure syntactic sugar (see lower.patlang's lower_match): the
# ["Match", scrutinee, arms, line] AST node it produces here carries its own
# line as the last element, same append-only convention as If/While/When.
make a function called at_case_stop takes toks, pos returns r
# Case-arm bodies stop at the next 'case', 'end', or EOF (never 'else'/
# 'elif' -- those belong to a nested if/while inside the arm body, not to
# match's own grammar, so at_block_stop's stop-set doesn't apply here).
let t = vec_get(toks, pos)
if t[0] == "EOF" then
return true
else
if (t[0] == "IDENT") and ((t[1] == "end") or (t[1] == "case")) then
return true
else
return false
end
end
end
make a function called parse_case_stmts_until takes toks, pos returns r
let stmts = []
let p = skip_nl(toks, pos)
let looping = true
while looping do
if at_case_stop(toks, p) then
let looping = false
else
let s = parse_stmt(toks, p)
let stmts = list_push(stmts, s[0])
let p = skip_nl(toks, s[1])
end
end
return [stmts, p]
end
# Collects a run of IDENT/NUM/OP("*")/OP("?") token text into a single glob
# string, stopping before 'then'/'when' (or NL/EOF). Handles both leading-
# wildcard globs (`*fred*`, pos already at the leading '*') and trailing-
# wildcard globs (`fred*`, pos at the leading identifier).
make a function called parse_glob_pattern takes toks, pos returns r
let glob_b = sb_new()
let p = pos
let looping = true
while looping do
let t = vec_get(toks, p)
if (t[0] == "IDENT") and ((t[1] == "then") or (t[1] == "when")) then
let looping = false
else
if (t[0] == "NL") or (t[0] == "EOF") then
let looping = false
else
if (t[0] == "OP") and ((t[1] == "*") or (t[1] == "?") or (t[1] == ".")) then
sb_push(glob_b, t[1])
let p = p + 1
else
if (t[0] == "IDENT") or (t[0] == "NUM") then
sb_push(glob_b, t[1])
let p = p + 1
else
let looping = false
end
end
end
end
end
return [["PGlob", sb_str(glob_b)], p]
end
make a function called parse_pattern takes toks, pos returns r
let t = vec_get(toks, pos)
if tok_is(t, "IDENT", "_") then
return [["PWild"], pos + 1]
else
if tok_is(t, "OP", "[") then
let p = skip_nl(toks, pos + 1)
if tok_is(vec_get(toks, p), "OP", "]") then
return [["PList", []], p + 1]
else
return parse_pattern_list_tail(toks, p, [])
end
else
if (t[0] == "OP") and ((t[1] == ">") or (t[1] == "<") or (t[1] == ">=") or (t[1] == "<=") or (t[1] == "==") or (t[1] == "!=")) then
let e = parse_expr(toks, pos + 1)
return [["PCmp", t[1], e[0]], e[1]]
else
if tok_is(t, "OP", "*") then
return parse_glob_pattern(toks, pos)
else
if t[0] == "NUM" then
return [["PLit", ["Num", t[1]]], pos + 1]
else
if t[0] == "STR" then
return [["PLit", ["Str", t[1]]], pos + 1]
else
if tok_is(t, "IDENT", "true") or tok_is(t, "IDENT", "false") then
return [["PLit", ["Bool", t[1]]], pos + 1]
else
if t[0] == "IDENT" then
let nxt = vec_get(toks, pos + 1)
if (nxt[0] == "OP") and ((nxt[1] == "*") or (nxt[1] == "?")) then
return parse_glob_pattern(toks, pos)
else
return [["PBind", t[1]], pos + 1]
end
else
return [["Err", "invalid pattern", t[2]], pos + 1]
end
end
end
end
end
end
end
end
end
make a function called parse_pattern_list_tail takes toks, pos, acc returns r
let pr = parse_pattern(toks, pos)
let acc2 = list_push(acc, pr[0])
let p = skip_nl(toks, pr[1])
let t = vec_get(toks, p)
if tok_is(t, "OP", ",") then
return parse_pattern_list_tail(toks, skip_nl(toks, p + 1), acc2)
else
if tok_is(t, "OP", "]") then
return [["PList", acc2], p + 1]
else
return [["Err", "expected ',' or ']' in list pattern", t[2]], p]
end
end
end
make a function called parse_case_arm takes toks, pos returns r
# pos points at 'case'. Returns [[pattern, guard_or_false, body], next_pos].
let pr = parse_pattern(toks, pos + 1)
let pattern = pr[0]
let p = skip_nl(toks, pr[1])
let t = vec_get(toks, p)
if tok_is(t, "IDENT", "when") then
let g = parse_expr(toks, p + 1)
let guard = g[0]
let p2 = skip_nl(toks, g[1])
let t2 = vec_get(toks, p2)
if tok_is(t2, "IDENT", "then") then
let body = parse_case_stmts_until(toks, p2 + 1)
return [[pattern, guard, body[0]], body[1]]
else
return [["Err", "expected then after when guard", t2[2]], p2]
end
else
if tok_is(t, "IDENT", "then") then
let body = parse_case_stmts_until(toks, p + 1)
return [[pattern, false, body[0]], body[1]]
else
return [["Err", "expected then (or when GUARD then) after case pattern", t[2]], p]
end
end
end
make a function called parse_match_arms takes toks, pos, acc returns r
let p = skip_nl(toks, pos)
let t = vec_get(toks, p)
if tok_is(t, "IDENT", "end") then
return [acc, p + 1]
else
if tok_is(t, "IDENT", "case") then
let arm = parse_case_arm(toks, p)
if arm[0][0] == "Err" then
return [arm[0], arm[1]]
else
return parse_match_arms(toks, skip_nl(toks, arm[1]), list_push(acc, arm[0]))
end
else
return [["Err", "expected case or end in match block", t[2]], p]
end
end
end
make a function called parse_match takes toks, pos returns r
# pos points at 'match'.
let match_line = vec_get(toks, pos)[2]
let se = parse_expr(toks, pos + 1)
let scrutinee = se[0]
let p = se[1]
let t = vec_get(toks, p)
if tok_is(t, "IDENT", "do") then
let arms = parse_match_arms(toks, p + 1, [])
if arms[0][0] == "Err" then
return [arms[0], arms[1]]
else
return [["Match", scrutinee, arms[0], match_line], arms[1]]
end
else
return [["Err", "expected do after match expression", t[2]], p]
end
end
make a function called parse_when takes toks, pos returns r
# pos points at 'when'; expect: when EVENT do body end, or when EVENT { body }
# `when_line` is appended as the node's LAST element (same append-only
# convention parse_if/parse_while already use for their own line
# numbers) so a later structural check (parse_check_when_placement)
# can report exactly where a misplaced `when` was written.
let when_line = vec_get(toks, pos)[2]
let nameTok = vec_get(toks, pos + 1)
if nameTok[0] == "IDENT" then
let ev = nameTok[1]
let p = pos + 2
if tok_is(vec_get(toks, p), "IDENT", "do") or tok_is(vec_get(toks, p), "IDENT", "begin") then
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
if tok_is(vec_get(toks, p2), "IDENT", "end") then
return [["When", ev, body[0], when_line], p2 + 1]
else
return [["Err", "expected end after when body", vec_get(toks, p2)[2]], p2]
end
else
if tok_is(vec_get(toks, p), "OP", "{") then
let body = parse_stmts_until(toks, p + 1)
let p2 = body[1]
if tok_is(vec_get(toks, p2), "OP", "}") then
return [["When", ev, body[0], when_line], p2 + 1]
else
return [["Err", "expected } after when body", vec_get(toks, p2)[2]], p2]
end
else
return [["Err", "expected do or { after when event", vec_get(toks, p)[2]], p]
end
end
else
return [["Err", "expected event name after when", vec_get(toks, pos + 1)[2]], pos + 1]
end
end
# ---- rule declarations: `rule Head(args) :- Body1, Body2.` / `rule Head(args).` ----
make a function called parse_rule_goal takes toks, pos returns r
# parses `name(args)` -> [[pred, [arg_expr_nodes]], next_pos]
let nameTok = vec_get(toks, pos)
let name = nameTok[1]
let lp = vec_get(toks, pos + 1)
if tok_is(lp, "OP", "(") then
let a = parse_args(toks, pos + 2)
return [[name, a[0]], a[1]]
else
return [["Err", "expected ( after rule goal name '" + name + "'", nameTok[2]], pos + 1]
end
end
make a function called parse_rule_decl takes toks, pos returns r
# pos points at the rule head's name (the 'rule' keyword itself was
# already consumed by the caller). Returns
# [["RuleDecl", head_pred, [head_args], [[pred,[args]], ...]], next_pos].
let head = parse_rule_goal(toks, pos)
let head_pred = head[0][0]
let head_args = head[0][1]
let p = head[1]
let t = vec_get(toks, p)
if tok_is(t, "OP", ":-") then
let p = p + 1
let body = []
let looping = true
while looping do
let p = skip_nl(toks, p)
let g = parse_rule_goal(toks, p)
let body = list_push(body, g[0])
let p = g[1]
let p = skip_nl(toks, p)
let t2 = vec_get(toks, p)
if tok_is(t2, "OP", ",") then
let p = p + 1
else
if tok_is(t2, "OP", ".") then
let p = p + 1
let looping = false
else
let looping = false
end
end
end
return [["RuleDecl", head_pred, head_args, body], p]
else
if tok_is(t, "OP", ".") then
return [["RuleDecl", head_pred, head_args, []], p + 1]
else
return [["Err", "expected ':-' or '.' after rule head", vec_get(toks, p)[2]], p]
end
end
end
# ---- goal declarations: `goal NAME { dep1(args), dep2(args) }` ----
make a function called parse_goal_decl takes toks, pos returns r
# pos points at the goal's name (the 'goal' keyword itself already
# consumed by the caller). Returns [["GoalDecl", name, deps], next_pos]
# where deps is [[pred, [arg_expr_nodes]], ...] -- each dependency term
# is parsed via parse_rule_goal, the exact same fact-term shape a rule
# body's goals use.
let nameTok = vec_get(toks, pos)
let name = nameTok[1]
let p = pos + 1
let openTok = vec_get(toks, p)
if tok_is(openTok, "OP", "{") == false then
return [["Err", "expected { after goal name", openTok[2]], p]
else
let p = skip_nl(toks, p + 1)
let t = vec_get(toks, p)
if tok_is(t, "OP", "}") then
return [["GoalDecl", name, []], p + 1]
else
let deps = []
let looping = true
while looping do
let p = skip_nl(toks, p)
let g = parse_rule_goal(toks, p)
let deps = list_push(deps, g[0])
let p = skip_nl(toks, g[1])
let t2 = vec_get(toks, p)
if tok_is(t2, "OP", ",") then
let p = skip_nl(toks, p + 1)
else
let looping = false
end
end
let t3 = vec_get(toks, p)
if tok_is(t3, "OP", "}") then
return [["GoalDecl", name, deps], p + 1]
else
return [["Err", "expected } to close goal block", vec_get(toks, p)[2]], p]
end
end
end
end
# Slice 1+2+3 of the classes/traits/inheritance feature (see the
# "synchronous-questing-metcalfe" plan) -- mirrors native parser.rs's
# Token::Class handling exactly. pos points at the class's name (the
# 'class' keyword itself already consumed by the caller). Returns
# [["ClassDecl", name, parent_or_empty_string, fields, methods, traits],
# next_pos] where fields is [[field_name, default_expr_node], ...],
# methods (Slice 2) is [[method_name, params, body], ...] -- a method is
# just an ordinary `make a function called NAME ... end`/`{ }` body
# declared inside the class block, reusing parse_function_def exactly --
# and traits (Slice 3) is [trait_name, ...] from a `traits A, B` line.
make a function called parse_class_decl takes toks, pos returns r
let nameTok = vec_get(toks, pos)
let name = nameTok[1]
let p = pos + 1
let parent = ""
let nt = vec_get(toks, p)
if (nt[0] == "IDENT") and (nt[1] == "inherits") then
let p = p + 1
let parentTok = vec_get(toks, p)
if parentTok[0] != "IDENT" then
return [["Err", "expected parent class name after 'inherits'", parentTok[2]], p]
else
let parent = parentTok[1]
let p = p + 1
end
end
let openTok = vec_get(toks, p)
# Slice 5: `{ ... }` or `do ... end` (`begin` accepted as a synonym for
# `do`, matching every other dual-form block in this parser -- see
# parse_while/parse_when above). `word_form` remembers which delimiter
# opened the block so the loop below knows whether to stop at `}` or at
# the word `end`.
if (tok_is(openTok, "OP", "{") == false) and (tok_is(openTok, "IDENT", "do") == false) and (tok_is(openTok, "IDENT", "begin") == false) then
return [["Err", "expected '{' or 'do'/'begin' after class header", openTok[2]], p]
else
let word_form = tok_is(openTok, "OP", "{") == false
let p = skip_nl(toks, p + 1)
let fields = []
let methods = []
let traits = []
let looping = true
while looping do
let t = vec_get(toks, p)
let at_block_end = false
if word_form then
if tok_is(t, "IDENT", "end") then
let at_block_end = true
end
else
if tok_is(t, "OP", "}") then
let at_block_end = true
end
end
if at_block_end then
let looping = false
else
if (t[0] == "IDENT") and (t[1] == "traits") then
let p = p + 1
let tlooping = true
while tlooping do
let tt = vec_get(toks, p)
if tt[0] != "IDENT" then
return [["Err", "expected trait name", tt[2]], p]
else
let traits = list_push(traits, tt[1])
let p = p + 1
if tok_is(vec_get(toks, p), "OP", ",") then
let p = p + 1
else
let tlooping = false
end
end
end
let p = skip_nl(toks, p)
else
if (t[0] == "IDENT") and (t[1] == "field") then
let fnameTok = vec_get(toks, p + 1)
if fnameTok[0] != "IDENT" then
return [["Err", "expected field name after 'field'", fnameTok[2]], p + 1]
else
let eqTok = vec_get(toks, p + 2)
if tok_is(eqTok, "OP", "=") == false then
return [["Err", "expected '=' after field name", eqTok[2]], p + 2]
else
let e = parse_expr(toks, p + 3)
let fields = list_push(fields, [fnameTok[1], e[0]])
let p = skip_nl(toks, e[1])
end
end
else
if (t[0] == "IDENT") and (t[1] == "make") then
let fr = parse_function_def(toks, p)
let fnode = fr[0]
if fnode[0] == "Err" then
return [fnode, fr[1]]
else
let methods = list_push(methods, [fnode[1], fnode[2], fnode[3]])
let p = skip_nl(toks, fr[1])
end
else
return [["Err", "expected 'field NAME = EXPR', 'make a function called NAME ... end', or 'traits A, B' inside a class block", t[2]], p]
end
end
end
end
end
return [["ClassDecl", name, parent, fields, methods, traits], p + 1]
end
end
make a function called parse_stmt_inner takes toks, pos returns r
let t = vec_get(toks, pos)
let ty = t[0]
let tx = t[1]
if tok_is(t, "OP", "{") then
# A bare `{ ... }` is only meaningful as a VALUE (parse_primary turns
# it into a zero-param closure literal, e.g. `let f = { ... }` -- a
# real function literal, not a leftover tolerance) -- never as its
# own free-standing statement. Rejecting it here is what makes
# dropping the statement-separator requirement safe: a bare block
# sitting right after something that could take a trailing-closure
# argument would otherwise be genuinely ambiguous between "two
# statements" and "one statement with sugar." Mirrors native
# parser.rs's parse_statement Token::BlockStart arm exactly.
return [["Err", "a bare '{ ... }' isn't a statement on its own -- it's a function literal as a VALUE (e.g. `let f = { ... }`, then call it with `f()`), not something to write standalone", t[2]], pos + 1]
else
if ty == "IDENT" then
if tx == "let" then
let mutTok = vec_get(toks, pos + 1)
let isMut = tok_is(mutTok, "IDENT", "mut")
let nameIdx = pos + 1
if isMut then
let nameIdx = pos + 2
end
let nameTok = vec_get(toks, nameIdx)
let name = nameTok[1]
let eqTok = vec_get(toks, nameIdx + 1)
if tok_is(eqTok, "OP", "=") then
let e = parse_expr(toks, nameIdx + 2)
return [["Let", name, e[0], false, isMut], e[1]]
else
return [["Err", "expected = after let name", vec_get(toks, nameIdx)[2]], nameIdx]
end
else
if tx == "if" then
return parse_if(toks, pos)
else
if tx == "while" then
return parse_while(toks, pos)
else
if tx == "match" then
return parse_match(toks, pos)
else
if (tx == "require") or (tx == "ensure") or (tx == "assert") then
let e = parse_expr(toks, pos + 1)
return [["Assert", tx, e[0]], e[1]]
else
if tx == "return" then
let e = parse_expr(toks, pos + 1)
return [["Return", e[0]], e[1]]
else
if tx == "make" then
return parse_function_def(toks, pos)
else
if tx == "fn" then
return parse_fn_def(toks, pos)
else
if tx == "when" then
return parse_when(toks, pos)
else
if (tx == "rule") and (tok_is(vec_get(toks, pos + 1), "OP", "(") == false) then
# Real declarative syntax: `rule Head(args) :- Body1, Body2.`
# (a rule) or `rule Head(args).` (a fact -- a rule with an
# empty body). The already-working call form `rule(...)`
# falls through to the bare-identifier-call path below,
# matching the native parser's own `if peek==LParen` branch.
return parse_rule_decl(toks, pos + 1)
else
if (tx == "goal") and (tok_is(vec_get(toks, pos + 1), "OP", "(") == false) then
# Real declarative syntax: `goal NAME { dep1(args), ... }`,
# mirroring native parser.rs's Token::Goal branch exactly.
return parse_goal_decl(toks, pos + 1)
else
if (tx == "class") and (tok_is(vec_get(toks, pos + 1), "OP", "(") == false) then
# Slice 1 of the classes/traits/inheritance feature (see
# the "synchronous-questing-metcalfe" plan): `class NAME
# [inherits PARENT] { field NAME = EXPR ... }`,
# mirroring native parser.rs's Token::Class branch
# exactly.
return parse_class_decl(toks, pos + 1)
else
if (tx == "pursue") or (tx == "activate") then
# Both are expressions (see parse_primary), routed through
# parse_expr/parse_primary at the statement level too --
# same pattern budgeted(...) uses just below.
let e = parse_expr(toks, pos)
return [["Expr", e[0]], e[1]]
else
if (tx == "budgeted") and tok_is(vec_get(toks, pos + 1), "OP", "(") then
# Bare statement usage (result discarded): route through
# parse_expr/parse_primary, which is where budgeted(...)
# is actually parsed -- parse_stmt's own bare-call fallback
# below has no block-body awareness and would otherwise
# mis-parse `budgeted(ms) { ... }`/`do ... end` as a
# malformed ordinary call.
let e = parse_expr(toks, pos)
return [["Expr", e[0]], e[1]]
else
# General fallback: parse a full expression (this already
# covers bare calls NAME(args), member/index chains
# obj.prop / obj.prop.sub / list[i], and full binary
# expressions like classify(x) == "adult" via the same
# parse_expr/parse_postfix/parse_primary path used
# everywhere else -- unlike native parser.rs, this
# self-hosted lexer/parser never treats a bare '=' as an
# equality operator (only '==' is recognized as a
# comparison, see the OP-token check just above this
# function), so parse_expr always stops cleanly BEFORE a
# pending assignment '=', with no risk of the ambiguity
# the native side has to guard against with its
# stop_trailing_block_for_condition flag. Previously this
# branch hand-rolled only 3 narrow shapes (bare call,
# bare reassignment, obj.prop[ = value]) and errored on
# anything else -- found via `pair[0]` (a bare index
# expression) and `classify(x) == "adult"` (a bare call
# followed by a comparison) both failing to parse as a
# closure body's implicit-return statement.
let e = parse_expr(toks, pos)
let node = e[0]
let p2 = e[1]
let eqTok = vec_get(toks, p2)
if tok_is(eqTok, "OP", "=") then
if node[0] == "Member" then
let vr = parse_expr(toks, p2 + 1)
return [["MemberAssign", node[1], node[2], vr[0]], vr[1]]
else
if node[0] == "Var" then
let vr = parse_expr(toks, p2 + 1)
return [["Let", node[1], vr[0], true, false], vr[1]]
else
return [["Err", "cannot assign to this expression", eqTok[2]], p2]
end
end
else
return [["Expr", node], p2]
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
else
return [["Err", "unexpected " + ty + " '" + tx + "'", vec_get(toks, pos)[2]], pos + 1]
end
end
end
# Thin line-tagging wrapper around parse_stmt_inner: the debugger's
# breakpoint/step machinery (self_hosting/lib/interp.patlang) needs a
# source line on Let/Expr/Return/Assert/MemberAssign nodes, which were
# never given one (If/While/When/Err already carry their own trailing
# line field). Deliberately an INCLUSION list, not "append to everything
# except the types that already have a line": Func nodes use
# `s.length > 4` in lower_program to detect an optional trailing
# return-hint element, so unconditionally appending a line there would
# silently corrupt that check on every ordinary function (s[4] would
# sometimes be the return hint, sometimes a line number). Any other node
# shape gains a shape-dependent meaning for a new trailing slot exactly
# the same way, so only touch the shapes actually needed here.
make a function called parse_stmt takes toks, pos returns r
let line = vec_get(toks, pos)[2]
let inner = parse_stmt_inner(toks, pos)
let node = inner[0]
let ty = node[0]
if (ty == "Let") or (ty == "Expr") or (ty == "Return") or (ty == "Assert") or (ty == "MemberAssign") then
return [list_push(node, line), inner[1]]
else
return inner
end
end
make a function called parse_program takes toks returns ast
let stmts = []
let pos = skip_nl(toks, 0)
let looping = true
while looping do
let t = vec_get(toks, pos)
if t[0] == "EOF" then
let looping = false
else
let r = parse_stmt(toks, pos)
let stmts = list_push(stmts, r[0])
let pos = skip_nl(toks, r[1])
end
end
return ["Program", stmts]
end
# ---- AST pretty printer ----
make a function called ast_list_to_str takes nodes returns s
# sb_new/sb_push/sb_str -- low real-world risk (this is a debug/error-
# message pretty printer, not part of the ordinary compile path), but
# fixed for consistency once the anti-pattern was found and a compiler
# warning added to catch it elsewhere.
let b = sb_new()
let i = 0
while i < nodes.length do
if i > 0 then
sb_push(b, "; ")
end
sb_push(b, ast_to_str(nodes[i]))
let i = i + 1
end
return sb_str(b)
end
# Joins a list of plain strings (e.g. closure param names), not AST nodes
make a function called ast_list_to_str_plain takes items returns s
let b = sb_new()
let i = 0
while i < items.length do
if i > 0 then
sb_push(b, ", ")
end
sb_push(b, items[i])
let i = i + 1
end
return sb_str(b)
end
# Renders a RuleDecl body: a list of [pred, arg_expr_nodes] pairs (not
# ast nodes with their own type tag, so this can't just call ast_to_str
# on each element the way ast_list_to_str does).
make a function called rule_body_to_str takes body returns s
let sb = sb_new()
sb_push(sb, "[")
let i = 0
let n = to_num(list_len(body))
while i < n do
if i > 0 then
sb_push(sb, ", ")
end
let g = body[i]
sb_push(sb, g[0] + "(" + ast_list_to_str(g[1]) + ")")
let i = i + 1
end
sb_push(sb, "]")
return sb_str(sb)
end
# GitHub #25 (Piece 1 investigation, contract_check under --x64):
# ast_to_str is a DEBUG/AST-inspector dump (its own "Var(b)"/"Num(0)"
# wrapper format is deliberate there -- playground_main.patlang uses it
# to show node TYPES explicitly). lower.patlang's require/ensure/assert
# lowering used it too, for the condition TEXT baked into a contract
# violation message ("contract violation: precondition failed in
# safe_divide(): ...") -- but that text is supposed to read like the
# original source (`b != 0`), matching rust-runtime/src/ir/lowering.rs's
# own expr_to_text exactly (the native --ir-run/--patc backends' own
# condition-stringifier), not an AST debug dump. This divergence was
# real and invisible for a long time: it only showed up once a real
# --x64 compile+run of a failing contract could happen at all (#98
# blocked that entirely before). A separate function, not a change to
# ast_to_str itself, so the debug-dump behavior elsewhere is unaffected.
make a function called ast_list_to_source_text takes nodes returns s
let b = sb_new()
let i = 0
while i < nodes.length do
if i > 0 then
sb_push(b, ", ")
end
sb_push(b, ast_to_source_text(nodes[i]))
let i = i + 1
end
return sb_str(b)
end
make a function called ast_to_source_text takes node returns s
let ty = node[0]
if ty == "Num" then
return node[1]
elif ty == "Str" then
return "\"" + node[1] + "\""
elif ty == "Bool" then
return node[1]
elif ty == "Var" then
return node[1]
elif ty == "Bin" then
return ast_to_source_text(node[2]) + " " + node[1] + " " + ast_to_source_text(node[3])
elif ty == "Un" then
return node[1] + ast_to_source_text(node[2])
elif ty == "Call" then
return node[1] + "(" + ast_list_to_source_text(node[2]) + ")"
elif ty == "List" then
return "[" + ast_list_to_source_text(node[1]) + "]"
elif ty == "Index" then
return ast_to_source_text(node[1]) + "[" + ast_to_source_text(node[2]) + "]"
elif ty == "Member" then
return ast_to_source_text(node[1]) + "." + node[2]
else
# Anything not covered above (Closure, MemberAssign, control-flow
# nodes -- not realistic inside a require/ensure/assert condition)
# falls back to the debug dump rather than crashing on an
# unrecognized shape.
return ast_to_str(node)
end
end
make a function called ast_to_str takes node returns s
let ty = node[0]
if ty == "Num" then
return "Num(" + node[1] + ")"
else
if ty == "Str" then
return "Str('" + node[1] + "')"
else
if ty == "Bool" then
return "Bool(" + node[1] + ")"
else
if ty == "Var" then
return "Var(" + node[1] + ")"
else
if ty == "Bin" then
return "(" + ast_to_str(node[2]) + " " + node[1] + " " + ast_to_str(node[3]) + ")"
else
if ty == "Un" then
return "(" + node[1] + " " + ast_to_str(node[2]) + ")"
else
if ty == "Call" then
return node[1] + "(" + ast_list_to_str(node[2]) + ")"
else
if ty == "Closure" then
return "|" + ast_list_to_str_plain(node[1]) + "| do " + ast_list_to_str(node[2]) + " end"
else
if ty == "List" then
return "[" + ast_list_to_str(node[1]) + "]"
else
if ty == "Index" then
return ast_to_str(node[1]) + "[" + ast_to_str(node[2]) + "]"
else
if ty == "Member" then
return ast_to_str(node[1]) + "." + node[2]
else
if ty == "Let" then
return "Let " + node[1] + " = " + ast_to_str(node[2])
else
if ty == "MemberAssign" then
return ast_to_str(node[1]) + "." + node[2] + " = " + ast_to_str(node[3])
else
if ty == "Expr" then
return ast_to_str(node[1])
else
if ty == "If" then
return "If " + ast_to_str(node[1]) + " Then {" + ast_list_to_str(node[2]) + "} Else {" + ast_list_to_str(node[3]) + "}"
else
if ty == "While" then
return "While " + ast_to_str(node[1]) + " {" + ast_list_to_str(node[2]) + "}"
else
if ty == "Func" then
return "Func " + node[1] + " {" + ast_list_to_str(node[3]) + "}"
else
if ty == "Return" then
return "Return " + ast_to_str(node[1])
else
if ty == "When" then
return "When " + node[1] + " {" + ast_list_to_str(node[2]) + "}"
else
if ty == "RuleDecl" then
return "RuleDecl(" + node[1] + ", " + ast_list_to_str(node[2]) + ", " + rule_body_to_str(node[3]) + ")"
else
if ty == "GoalDecl" then
return "GoalDecl(" + node[1] + ", " + rule_body_to_str(node[2]) + ")"
else
if ty == "Assert" then
return node[1] + " " + ast_to_str(node[2])
else
if ty == "Budgeted" then
return "budgeted(" + ast_to_str(node[1]) + ", " + ast_to_str(node[2]) + ") {" + ast_list_to_str(node[3]) + "}"
else
if ty == "Err" then
# Markers, not a single "@" -- the message can itself
# echo arbitrary token text (e.g. "unexpected UNK '@'"
# for a literal '@' in the source), so a single-char
# delimiter is not safe to search for. "@@ENDERR@@"
# gives parse_find_error an unambiguous message-end
# boundary regardless of what the message contains.
#
# Checking `ty == "Err"` explicitly here (not just "any
# unrecognized tag with >=3 elements", which is what this
# used to do) matters for a real reason, not just
# precision for its own sake: the OLD version produced a
# genuine false positive the moment this file's own
# source was compiled by patc1.exe and handed to itself
# as input (e.g. any selftest that `include`s lib/
# parser.patlang) -- this very file's source literally
# CONTAINS the string "@@ERR@@" as a string-literal value
# (right here), so a generic "any unknown 3+-field tag"
# fallback would render THIS Str node's own content and
# parse_find_error would mistake it for a real error, on
# perfectly valid, successfully-parsed source. Confirmed
# directly: self_hosting/regex_dsl_selftest.patlang
# (pre-existing, not new) failed to compile via patc1.exe
# with exactly this symptom before this fix.
return "@@ERR@@" + ("" + node[2]) + "@@" + node[1] + "@@ENDERR@@"
else
# A genuinely unrecognized tag (not "Err") -- this
# shouldn't normally happen given the catalog above, but
# stays a safety net for a future missing case (the way
# "RuleDecl" once was, before it got its own real case).
# Deliberately does NOT contain "@@ERR@@"/"@@ENDERR@@" or
# any other text parse_find_error searches for, so an
# unrecognized tag renders as inert, harmless text instead
# of being silently misidentified as a parse error.
return "Unrecognized(" + ty + ")"
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
end
# ---- compiler error reasoning: report *and* suggest a fix, not just
# fail silently. This is what closes the long-standing gap where
# patc1_main.patlang (the actual self-hosted compiler driver) had NO
# error handling at all -- a syntax error just silently produced an
# ["Err", ...] node buried somewhere in the AST, which lower_program
# either choked on unpredictably or silently dropped (the "silent
# statement dropping" Stage 0 gotcha), with zero indication to the user
# that anything went wrong. All PatLang-level: no rustc/Rust involved
# anywhere in this reasoning, and none ever should be -- the project's
# own direction is to remove that dependency entirely, not build more
# tooling around its error output. ----
# The first Err node found anywhere in a parsed program's statement
# list -- top-level, or nested inside a statement's own expressions
# (e.g. `let z = (((` produces a top-level Let statement whose VALUE is
# the malformed node, not a top-level Err itself).
#
# This is a genuine typed walk over the real AST (checking each node's
# OWN tag field directly), not text-rendering-and-scanning. An earlier
# version rendered each statement with ast_to_str and searched the
# rendered text for an "@@ERR@@...@@ENDERR@@" marker pair -- reusing
# self_hosting/lib/repl_core.patlang's original technique for the
# message alone, extended to also recover the line number. That
# approach had a real, confirmed bug: ast_to_str renders a plain Str
# node's OWN VALUE verbatim inline, so a string literal whose CONTENT
# happened to equal the marker text produced a false positive --
# exactly what happens the moment THIS FILE's own source (which
# legitimately contains "@@ERR@@" as a string literal, since that's the
# code that used to construct the marker) is compiled by patc1.exe and
# handed to itself as input, e.g. any selftest that `include`s this
# library. Confirmed directly: self_hosting/regex_dsl_selftest.patlang
# (pre-existing, unrelated to this session) failed to compile via
# patc1.exe with exactly this symptom. A typed walk checking `node[0] ==
# "Err"` directly is immune to this whole class of bug by construction
# -- it can never confuse a node's DATA (a string's contents) with its
# STRUCTURE (its tag), because it never turns the tree into text at all.
#
# Mirrors ast_to_str's own per-tag field layout exactly (read directly
# from that function, not guessed) so this stays a real inverse of it,
# not an approximation. Returns ["", -1] if no Err node is found
# anywhere.
make a function called parse_find_error takes stmts returns err
let placement = parse_check_when_placement(stmts, 0)
if placement[0] != "" then
return placement
end
return pf_walk_list(stmts)
end
make a function called pf_walk_list takes nodes returns err
let i = 0
let n = to_num(list_len(nodes))
while i < n do
let r = pf_walk_node(nodes[i])
if r[0] != "" then
return r
end
let i = i + 1
end
return ["", -1]
end
make a function called pf_walk_pattern takes pattern returns err
let tag = pattern[0]
if tag == "Err" then
return [pattern[1], pattern[2]]
end
if tag == "PCmp" then
return pf_walk_node(pattern[2])
end
if tag == "PList" then
let subs = pattern[1]
let i = 0
let n = to_num(list_len(subs))
while i < n do
let r = pf_walk_pattern(subs[i])
if r[0] != "" then return r end
let i = i + 1
end
return ["", -1]
end
# PWild/PBind/PLit/PGlob carry no sub-expressions that could hold an Err.
return ["", -1]
end
make a function called pf_walk_node takes node returns err
let ty = node[0]
if ty == "Err" then
return [node[1], node[2]]
end
if (ty == "Num") or (ty == "Str") or (ty == "Bool") or (ty == "Var") then
return ["", -1]
end
if ty == "Bin" then
let r = pf_walk_node(node[2])
if r[0] != "" then return r end
return pf_walk_node(node[3])
end
if ty == "Un" then
return pf_walk_node(node[2])
end
if (ty == "Call") or (ty == "Closure") then
return pf_walk_list(node[2])
end
if ty == "List" then
return pf_walk_list(node[1])
end
if ty == "Index" then
let r = pf_walk_node(node[1])
if r[0] != "" then return r end
return pf_walk_node(node[2])
end
if ty == "Member" then
return pf_walk_node(node[1])
end
if ty == "Let" then
return pf_walk_node(node[2])
end
if ty == "MemberAssign" then
let r = pf_walk_node(node[1])
if r[0] != "" then return r end
return pf_walk_node(node[3])
end
if (ty == "Expr") or (ty == "Return") then
return pf_walk_node(node[1])
end
if ty == "If" then
let r = pf_walk_node(node[1])
if r[0] != "" then return r end
let r2 = pf_walk_list(node[2])
if r2[0] != "" then return r2 end
return pf_walk_list(node[3])
end
if ty == "While" then
let r = pf_walk_node(node[1])
if r[0] != "" then return r end
return pf_walk_list(node[2])
end
if ty == "Match" then
let r = pf_walk_node(node[1])
if r[0] != "" then return r end
let arms = node[2]
let i = 0
let n = to_num(list_len(arms))
while i < n do
let arm = arms[i]
let pr = pf_walk_pattern(arm[0])
if pr[0] != "" then return pr end
if arm[1] != false then
let gr = pf_walk_node(arm[1])
if gr[0] != "" then return gr end
end
let br = pf_walk_list(arm[2])
if br[0] != "" then return br end
let i = i + 1
end
return ["", -1]
end
if (ty == "Func") or (ty == "When") then
return pf_walk_list(node[2])
end
if ty == "Assert" then
return pf_walk_node(node[2])
end
if ty == "Budgeted" then
let r = pf_walk_node(node[1])
if r[0] != "" then return r end
let r2 = pf_walk_node(node[2])
if r2[0] != "" then return r2 end
return pf_walk_list(node[3])
end
if ty == "RuleDecl" then
let r = pf_walk_list(node[2])
if r[0] != "" then return r end
let body = node[3]
let i = 0
let n = to_num(list_len(body))
while i < n do
let r2 = pf_walk_list(body[i][1])
if r2[0] != "" then return r2 end
let i = i + 1
end
return ["", -1]
end
if ty == "GoalDecl" then
let deps = node[2]
let i = 0
let n = to_num(list_len(deps))
while i < n do
let r = pf_walk_list(deps[i][1])
if r[0] != "" then return r end
let i = i + 1
end
return ["", -1]
end
# Unrecognized tag, or a leaf shape with nothing further to walk (e.g.
# "Bool"/"Num"/"Var" already handled above) -- nothing more to check.
return ["", -1]
end
# `when` is a genuine top-level-only declaration -- handlers are
# collected once, statically, when the program is lowered (rust-runtime/
# src/ir/lowering.rs's own top-level scan for Stmt::When; the self-
# hosted lower.patlang mirrors this). A `when` written anywhere else
# (nested inside if/while/a function body) used to compile with NO
# error at all and simply never fire, ever, silently -- confirmed
# directly while building self_hosting/examples/microservices_demo.
# patlang: `if role == "greeter" then when greet do ... end end`
# produced a worker that received every signal and answered NONE of
# them, no diagnostic anywhere pointing at why.
#
# This is deliberate, not a limitation to work around: PatLang's
# capability-discovery system (signal_discovery.patlang) treats a
# service's *advertised* actions as a fixed description of what it
# responds to for the life of the process. If handler registration
# could vary at runtime, an honest self-description could go stale on
# its own without a new announcement ever being made -- see that
# module's own header for the existing, separate "self-reported, not
# verified" trust caveat, which this would have quietly compounded.
# Global-only handlers were chosen deliberately over allowing
# conditional registration for that reason, not just because it's
# simpler to implement (though it is) -- so this is enforced as a real
# parse error now, not merely documented as a gotcha to remember.
#
# Walks the full statement tree (top-level, then recursively into every
# If/While/Func body, each one level deeper) looking for a "When" node
# at depth > 0. Returns [message, line] on the FIRST one found, or
# ["", -1] if every `when` in the program is genuinely at top level.
make a function called parse_check_when_placement takes stmts, depth returns err
let i = 0
let n = to_num(list_len(stmts))
while i < n do
let stmt = stmts[i]
let tag = stmt[0]
if tag == "When" then
if depth > 0 then
return ["'when' can only be declared at the top level of a program -- it cannot be nested inside if/while blocks or function bodies, since handlers are registered once, statically, not conditionally at runtime", stmt[3]]
end
end
if tag == "If" then
let r1 = parse_check_when_placement(stmt[2], depth + 1)
if r1[0] != "" then
return r1
end
let r2 = parse_check_when_placement(stmt[3], depth + 1)
if r2[0] != "" then
return r2
end
end
if tag == "While" then
let r = parse_check_when_placement(stmt[2], depth + 1)
if r[0] != "" then
return r
end
end
if tag == "Match" then
let arms = stmt[2]
let ai = 0
let an = to_num(list_len(arms))
while ai < an do
let r = parse_check_when_placement(arms[ai][2], depth + 1)
if r[0] != "" then
return r
end
let ai = ai + 1
end
end
if tag == "Func" then
let r = parse_check_when_placement(stmt[3], depth + 1)
if r[0] != "" then
return r
end
end
let i = i + 1
end
return ["", -1]
end
make a function called pf_find_substr takes hay, needle returns idx
let hn = hay.length
let nn = needle.length
if nn == 0 then
return 0
end
let i = 0
while i <= hn - nn do
if substr(hay, i, nn) == needle then
return i
end
let i = i + 1
end
return -1
end
# The text of source line `line` (1-indexed), with no trailing newline.
make a function called source_line_text takes source, line returns text
let n = source.length
let cur_line = 1
let start = 0
let i = 0
while (i < n) and (cur_line < line) do
if char_code(source, i) == 10 then
let cur_line = cur_line + 1
let start = i + 1
end
let i = i + 1
end
let end_pos = start
while (end_pos < n) and (char_code(source, end_pos) != 10) do
let end_pos = end_pos + 1
end
return substr(source, start, end_pos - start)
end
# Reasoning over the CLOSED, known set of parser error messages (see the
# ~30 return sites throughout this file) -- a direct, tailored suggestion
# per message shape, not a generic catch-all, matching the same "specific
# beats generic" discipline self_hosting/lib/synthesis_lgg.patlang's
# synth5_format_diagnosis already established for induction-engine
# diagnoses. Unrecognized messages (a future error site added without a
# matching case here) fall through to a still-useful generic line.
make a function called pf_contains takes hay, needle returns found
return pf_find_substr(hay, needle) >= 0
end
make a function called suggest_parse_fix takes message returns suggestion
if pf_contains(message, "'when' can only be declared at the top level") then
return "Move this 'when' block out to the top level of the program. If you need different behavior for different cases, register the handler unconditionally and put the conditional logic INSIDE its body instead (check whatever condition you need as the handler's first statement)."
end
if pf_contains(message, "unexpected") then
return "This token wasn't expected here -- look for a missing keyword, operator, or unbalanced bracket/parenthesis just before it."
end
if pf_contains(message, "expected end after") then
return "This block needs a matching 'end' -- check that every 'do'/'then'/block-opening keyword earlier has one, and that none was consumed by a nested block closing early."
end
if pf_contains(message, "else, elif, or end") then
return "An 'if' block needs to close with 'end', or continue with 'elif <cond> then'/'else' -- check what follows the last statement in this if-block."
end
if pf_contains(message, "} after") then
return "This block needs a matching '}' -- check for a missing closing brace, or an extra '{' earlier that consumed the one meant for this block."
end
if pf_contains(message, "expected )") then
return "A '(' opened earlier is missing its matching ')' -- check argument lists and parenthesized expressions for a missing closing paren, or a stray extra argument."
end
if pf_contains(message, "expected ]") then
return "A '[' opened earlier is missing its matching ']' -- check list literals and index expressions for a missing closing bracket."
end
if pf_contains(message, "expected = after let name") then
return "A 'let' declaration needs '= <expression>' right after the variable name."
end
if pf_contains(message, "'main' is a reserved") then
return "Rename this function -- 'main' collides with the native-compiled backend's own program entry point. Pick any other name and call it explicitly from top-level code instead."
end
if pf_contains(message, "function name") or pf_contains(message, "'function' after make") then
return "Function declarations need the exact form 'make a function called NAME takes ... returns ... ... end' -- check the keywords right after 'make'."
end
if pf_contains(message, "event name") or pf_contains(message, "after when event") then
return "'when' needs an event name identifier right after it, e.g. 'when some_event do ... end'."
end
if pf_contains(message, ":-") then
return "A rule head needs either ':-' followed by its body, or a '.' for a fact with no body."
end
if pf_contains(message, "name after .") then
return "Member access ('.') needs a plain identifier right after it, e.g. 'obj.field'."
end
return "Check the syntax immediately around this point against the construct being parsed."
end
# The full human-readable diagnostic: message, source location with the
# actual offending line shown, and a suggested fix -- exactly the "report
# on it, but also suggest potential fixes" the compiler-error-reasoning
# backlog item asked for, scoped to PatLang's own errors only.
make a function called format_parse_error takes message, line, source returns text
let line_text = source_line_text(source, line)
if line_text == "" then
let line_text = "(reached end of file -- nothing follows on this line)"
end
let suggestion = suggest_parse_fix(message)
return "Parse error at line " + ("" + line) + ":\n " + line_text + "\n " + message + "\n Suggestion: " + suggestion
end
# =============================================================================
# Stage 1 self-hosted lowerer (Stage 0 compilable subset).
# Walks the list-shaped AST from lib/parser.patlang and emits list-shaped IR
# instructions. The host's compile_ir only decodes this IR and runs codegen —
# lexing, parsing, lowering, AND (via lib/codegen.patlang) code generation all
# happen in PatLang.
#
# IR shape:
# ["ProgramIR", entry, [functions], [events]]
# ["FuncIR", name, [params], [instrs], [lines]]
# ["EventIR", event, handler_name]
#
# lines: a SPARSE, statement-granularity debug table, one [pc, source_line]
# pair per Let/Expr/Return/Assert/MemberAssign statement lowered into this
# function's own instrs (not every instruction -- see parser.patlang's
# parse_stmt wrapper for which node shapes carry a line at all, and why
# Func/If/While/When/Err were deliberately left out of that set). Built for
# self_hosting/lib/interp.patlang's debugger: enough to map a paused pc back
# to "which statement", not full per-instruction provenance.
# Instructions:
# ["Const", "num"|"str"|"bool", text] ["Load", name] ["Store", name]
# ["Bin", op] ["Un", op] ["Jump", n] ["JumpIfFalse", n]
# ["CallHost", name, argc] ["Call", name, argc]
# ["MakeClosure", func_name, [captured_names]] ["CallValue", argc]
# ["BuildList", n] ["Return"]
#
# Closures: |params| do body end (Stage 1 has no brace-delimited blocks
# anywhere, so closures use the same do/end convention as while-loops rather
# than Stage 0's |params| { body }). A closure captures its ENTIRE enclosing
# locals list (over-capture, not precise free-variable analysis) as leading
# hidden parameters of a synthesized function — simpler than exact capture
# and still correct: the flat per-call locals map means a closure's own
# `let` of a same-named variable just overwrites the pre-bound captured
# value, exactly matching intended shadowing.
# =============================================================================
make a function called contains_str takes xs, s returns r
let i = 0
while i < xs.length do
if xs[i] == s then
return true
end
let i = i + 1
end
return false
end
make a function called vec_contains takes v, s returns r
let i = 0
let n = vec_len(v)
while i < n do
if vec_get(v, i) == s then
return true
end
let i = i + 1
end
return false
end
# Finds the mutability flag for the MOST RECENT declaration of `name` in
# `locals`/`mutables` (index-parallel vecs) -- a re-`let` shadows, so the
# latest declaration's flag is the one that governs a later reassignment.
make a function called find_mut_flag takes locals, mutables, name returns r
let n = vec_len(locals)
let i = n - 1
let found = false
let result = false
while (i >= 0) and (not found) do
if vec_get(locals, i) == name then
let found = true
let result = vec_get(mutables, i)
end
let i = i - 1
end
return result
end
make a function called next_closure_name returns name
let n = get("__vars", "__closure_seq")
if not n then
let n = 0
end
set_var("__closure_seq", n + 1)
return "__closure_" + n
end
make a function called next_budgeted_name returns name
let n = get("__vars", "__budgeted_seq")
if not n then
let n = 0
end
set_var("__budgeted_seq", n + 1)
return "__budgeted_" + n
end
make a function called next_match_name returns name
let n = get("__vars", "__match_seq")
if not n then
let n = 0
end
set_var("__match_seq", n + 1)
return "__match_" + n
end
make a function called next_activate_name returns name
let n = get("__vars", "__activate_seq")
if not n then
let n = 0
end
set_var("__activate_seq", n + 1)
return "__act_" + n
end
# Nesting-depth counter (not a bool, so a budgeted block nested inside
# another still counts as "inside") tracking whether the statement/body
# currently being lowered is lexically inside a budgeted(...) { ... } block
# -- while-loop back-edges check this to decide whether to inject a
# budget_check() call. Stored via the same object-store convention as
# __closure_seq/__budgeted_seq since lower_*'s functions are plain
# functions, not methods on a stateful struct (unlike Stage 0's Lowerer).
make a function called budgeted_depth_get returns n
let n = get("__vars", "__budgeted_depth")
if not n then
return 0
end
return n
end
make a function called budgeted_depth_inc returns done
set_var("__budgeted_depth", budgeted_depth_get() + 1)
return true
end
make a function called budgeted_depth_dec returns done
set_var("__budgeted_depth", budgeted_depth_get() - 1)
return true
end
# lower_expr(node, code, fns, locals, pending, lines) -> code
# fns: list of top-level function names (static call vs host call)
# locals: vec of names currently bound in the enclosing scope (for
# disambiguating a call through a closure-valued variable, and
# for over-capturing into new closures)
# pending: vec of synthesized ["FuncIR", ...] nodes from closure literals,
# merged into the program's function list once lowering finishes
make a function called lower_expr takes node, code, fns, locals, pending, lines returns out
let ty = node[0]
if ty == "Num" then
vec_push(code, ["Const", "num", node[1]])
return code
else
if ty == "Str" then
vec_push(code, ["Const", "str", node[1]])
return code
else
if ty == "Bool" then
vec_push(code, ["Const", "bool", node[1]])
return code
else
if ty == "Var" then
# `unit` is the language's own literal for the Unit/void value
# (a program can write it directly as an expression, e.g.
# `return unit` -- see self_hosting/lib/interp.patlang's
# interp_const), so a bare `unit` that ISN'T a bound variable
# lowers to a Unit constant rather than a variable load. Found
# during the all-self-hosted fixpoint self-compile: without
# this, codegen_x64.patlang's own undefined-variable check
# correctly rejected the resulting unresolved `Load "unit"`.
#
# The `vec_contains(locals, ...)` guard is load-bearing, and
# NOT what rust-runtime/src/ir/lowering.rs's own Expr::
# Identifier arm does -- it special-cases the name
# unconditionally. Matching that unconditionally here is a
# real regression, because `unit` is also just an ordinary,
# legal variable name that this very codebase uses: x64_
# compile_unit.patlang's own x64_make_compile_unit declares
# `returns unit`, does `let unit = new("CompileUnit", ...)`,
# and ends `return unit`. Lowering that final read as a Unit
# constant makes the function silently return void instead of
# the object it just built -- every per-function compile unit
# then came back empty, and every --x64 build died with a
# confusing "LINK ERROR: undefined symbol '<first function>'"
# far downstream of the actual cause. A bound name always wins
# over the literal.
if (node[1] == "unit") and (not vec_contains(locals, "unit")) then
vec_push(code, ["Const", "unit", ""])
return code
end
vec_push(code, ["Load", node[1]])
return code
else
if ty == "Bin" then
if node[1] == "and" then
# short-circuit: false when lhs falsey, else truthiness of rhs
let code = lower_expr(node[2], code, fns, locals, pending, lines)
vec_push(code, ["Un", "not"])
let jif = vec_len(code)
vec_push(code, ["JumpIfFalse", 0])
vec_push(code, ["Const", "bool", "false"])
let jmp = vec_len(code)
vec_push(code, ["Jump", 0])
vec_set(code, jif, ["JumpIfFalse", vec_len(code)])
let code = lower_expr(node[3], code, fns, locals, pending, lines)
vec_push(code, ["Un", "not"])
vec_push(code, ["Un", "not"])
vec_set(code, jmp, ["Jump", vec_len(code)])
return code
else
if node[1] == "or" then
# short-circuit: true when lhs truthy, else truthiness of rhs
let code = lower_expr(node[2], code, fns, locals, pending, lines)
vec_push(code, ["Un", "not"])
let jif = vec_len(code)
vec_push(code, ["JumpIfFalse", 0])
let code = lower_expr(node[3], code, fns, locals, pending, lines)
vec_push(code, ["Un", "not"])
vec_push(code, ["Un", "not"])
let jmp = vec_len(code)
vec_push(code, ["Jump", 0])
vec_set(code, jif, ["JumpIfFalse", vec_len(code)])
vec_push(code, ["Const", "bool", "true"])
vec_set(code, jmp, ["Jump", vec_len(code)])
return code
else
let code = lower_expr(node[2], code, fns, locals, pending, lines)
let code = lower_expr(node[3], code, fns, locals, pending, lines)
vec_push(code, ["Bin", node[1]])
return code
end
end
else
if ty == "Un" then
let code = lower_expr(node[2], code, fns, locals, pending, lines)
vec_push(code, ["Un", node[1]])
return code
else
if ty == "Call" then
let callee = node[1]
let args = node[2]
if contains_str(fns, callee) then
# GitHub #102/#50: automatic box-on-crossing for a
# genuinely ambiguous call argument is now done as an
# AST REWRITE PASS (x64_rewrite_stmts_boxing, run once
# by lower_program before this function ever sees
# stmts) rather than here -- that pass can compute
# each function's own float-tainted-variable context
# once and thread it through naturally, which this
# per-Call codegen site has no way to do (lower_expr
# carries no "which function/scope am I in" parameter).
# By the time control reaches here, any needed
# rt_box_float wrapping is already a literal Call node
# in `args[i]` -- nothing extra to do.
let i = 0
while i < args.length do
let code = lower_expr(args[i], code, fns, locals, pending, lines)
let i = i + 1
end
vec_push(code, ["Call", callee, args.length])
return code
else
if vec_contains(locals, callee) then
# dynamic call through a local variable holding a closure
vec_push(code, ["Load", callee])
let i = 0
while i < args.length do
let code = lower_expr(args[i], code, fns, locals, pending, lines)
let i = i + 1
end
vec_push(code, ["CallValue", args.length])
return code
else
let i = 0
while i < args.length do
let code = lower_expr(args[i], code, fns, locals, pending, lines)
let i = i + 1
end
vec_push(code, ["CallHost", callee, args.length])
return code
end
end
else
if ty == "List" then
let items = node[1]
let i = 0
while i < items.length do
let code = lower_expr(items[i], code, fns, locals, pending, lines)
let i = i + 1
end
vec_push(code, ["BuildList", items.length])
return code
else
if ty == "Index" then
let code = lower_expr(node[1], code, fns, locals, pending, lines)
let code = lower_expr(node[2], code, fns, locals, pending, lines)
vec_push(code, ["CallHost", "list_get", 2])
return code
else
if ty == "Member" then
let code = lower_expr(node[1], code, fns, locals, pending, lines)
if (node[2] == "length") or (node[2] == "len") then
vec_push(code, ["CallHost", "len", 1])
return code
else
vec_push(code, ["Const", "str", node[2]])
vec_push(code, ["CallHost", "get", 2])
return code
end
else
if ty == "MethodCall" then
# obj.method(args) -- send(obj, "method", ...args),
# the same host call obj.prop = value already uses
# on the write side (see Stmt::MemberAssign below).
let object = node[1]
let method = node[2]
let args = node[3]
let code = lower_expr(object, code, fns, locals, pending, lines)
vec_push(code, ["Const", "str", method])
let i = 0
while i < args.length do
let code = lower_expr(args[i], code, fns, locals, pending, lines)
let i = i + 1
end
vec_push(code, ["CallHost", "send", 2 + args.length])
return code
else
if ty == "Closure" then
return lower_closure_literal(node[1], node[2], code, fns, locals, pending, lines)
else
if ty == "Budgeted" then
return lower_budgeted_literal(node[1], node[2], node[3], code, fns, locals, pending, lines)
else
# unknown expression: lower as unit-ish empty string
vec_push(code, ["Const", "str", ""])
return code
end
end
end
end
end
end
end
end
end
end
end
end
end
end
# Lowers a closure literal: synthesizes a Function (params = the enclosing
# scope's current locals ++ the closure's own params) queued in `pending`,
# and emits the Load.../MakeClosure sequence at the creation site into `code`.
make a function called lower_closure_literal takes params, body, code, fns, locals, pending, lines returns out
let captured = vec_to_list(locals)
let func_name = next_closure_name()
let all_params = []
let i = 0
while i < captured.length do
let all_params = list_push(all_params, captured[i])
let i = i + 1
end
let i = 0
while i < params.length do
let all_params = list_push(all_params, params[i])
let i = i + 1
end
let inner_locals = vec_new()
let inner_mutables = vec_new()
let i = 0
while i < all_params.length do
vec_push(inner_locals, all_params[i])
vec_push(inner_mutables, true)
let i = i + 1
end
# A closure's body is its own function with its own pc-space, so it gets
# its own fresh lines table -- the threaded-in `lines` param belongs to
# the OUTER code this closure literal is embedded in, and recording the
# inner body's statements into it would attach the wrong function's pcs
# to those source lines.
let inner_lines = vec_new()
let inner_code = lower_block(body, vec_new(), fns, func_name, inner_locals, inner_mutables, pending, inner_lines, true)
vec_push(inner_code, ["Return"])
vec_push(pending, ["FuncIR", func_name, all_params, vec_to_list(inner_code), vec_to_list(inner_lines)])
# At the creation site: push captured values in the SAME order used for
# all_params's leading section, then bundle them.
let i = 0
while i < captured.length do
vec_push(code, ["Load", captured[i]])
let i = i + 1
end
vec_push(code, ["MakeClosure", func_name, captured])
return code
end
# Lowers budgeted(ms[, existing]) { body }: synthesizes a function (like a
# closure) run via an implicit fiber. Unlike closures, captured locals are
# bundled into a SINGLE List argument rather than passed as leading params --
# fiber_new only ever passes one initial argument to the spawned function --
# and the synthesized function's own prologue unpacks them again via
# list_get/Store. `existing` defaults to Bool(false) at parse time when
# omitted from source (see parser.patlang), which budgeted_run treats the
# same as Unit: "no handle yet, start a new fiber".
make a function called lower_budgeted_literal takes msNode, existingNode, body, code, fns, locals, pending, lines returns out
let captured = vec_to_list(locals)
let func_name = next_budgeted_name()
let inner_locals = vec_new()
let inner_mutables = vec_new()
vec_push(inner_locals, "__captured")
vec_push(inner_mutables, true)
let i = 0
while i < captured.length do
vec_push(inner_locals, captured[i])
vec_push(inner_mutables, true)
let i = i + 1
end
let inner_code = vec_new()
let i = 0
while i < captured.length do
vec_push(inner_code, ["Load", "__captured"])
vec_push(inner_code, ["Const", "num", "" + i])
vec_push(inner_code, ["CallHost", "list_get", 2])
vec_push(inner_code, ["Store", captured[i]])
let i = i + 1
end
# Same reasoning as lower_closure_literal: this is a separate function
# with its own pc-space, so it needs its own fresh lines table rather
# than the caller's.
let inner_lines = vec_new()
budgeted_depth_inc()
let inner_code = lower_block(body, inner_code, fns, func_name, inner_locals, inner_mutables, pending, inner_lines, true)
budgeted_depth_dec()
vec_push(inner_code, ["Return"])
vec_push(pending, ["FuncIR", func_name, ["__captured"], vec_to_list(inner_code), vec_to_list(inner_lines)])
# budgeted_run(ms, func_name, captured_list, existing)
let code = lower_expr(msNode, code, fns, locals, pending, lines)
vec_push(code, ["Const", "str", func_name])
let i = 0
while i < captured.length do
vec_push(code, ["Load", captured[i]])
let i = i + 1
end
vec_push(code, ["BuildList", captured.length])
let code = lower_expr(existingNode, code, fns, locals, pending, lines)
vec_push(code, ["CallHost", "budgeted_run", 4])
return code
end
# Lowers `activate(PLAN)` (the desugared form of `activate PLAN`, see
# parser.patlang's parse_primary) -- mirrors ir/lowering.rs's
# lower_activate exactly: host functions can't call back into the
# interpreter to invoke a closure, so this is synthesized as ordinary
# control-flow AST (Let/While/If/Call) fed through the normal
# lower_block/lower_stmt path, rather than a new IR instruction. Each
# plan step's bound closure receives its bound argument values as a
# single List (not positionally splatted -- a plan step's real arity is
# only known once the planner has run), and the loop stops immediately,
# reporting failure, the moment one returns false.
make a function called lower_activate_stmt takes plan_node, code, fns, fname, locals, mutables, pending, lines returns out
let n = next_activate_name()
let plan_var = "__act_plan_" + n
let i_var = "__act_i_" + n
let ok_var = "__act_ok_" + n
let label_var = "__act_label_" + n
let fn_var = "__act_fn_" + n
let args_var = "__act_args_" + n
let r_var = "__act_r_" + n
let synth = [
["Let", plan_var, plan_node, false, false],
["Let", i_var, ["Num", "0"], false, true],
["Let", ok_var, ["Bool", "true"], false, true],
["While",
["Bin", "and",
["Var", ok_var],
["Bin", "<", ["Var", i_var], ["Call", "to_num", [["Call", "list_len", [["Var", plan_var]]]]]]
],
[
["Let", label_var, ["Index", ["Var", plan_var], ["Var", i_var]], false, false],
["Let", fn_var, ["Call", "action_lookup", [["Call", "action_base_name", [["Var", label_var]]]]], false, false],
["Let", args_var, ["Call", "action_label_args", [["Var", label_var]]], false, false],
["Let", r_var, ["Call", fn_var, [["Var", args_var]]], false, false],
["If", ["Bin", "==", ["Var", r_var], ["Bool", "false"]],
[["Let", ok_var, ["Bool", "false"], true, false]],
[]
],
["Let", i_var, ["Bin", "+", ["Var", i_var], ["Num", "1"]], true, false]
]
]
]
let code = lower_block(synth, code, fns, fname, locals, mutables, pending, lines, false)
return lower_expr(["Var", ok_var], code, fns, locals, pending, lines)
end
# Render a rule head/body arg as the compile-time string token rule_add
# expects: a bare identifier/number renders as its own text, but a string
# literal renders as its raw content, not source-quoted -- mirrors
# ir/lowering.rs's own rule_arg_text.
make a function called rule_arg_str takes node returns s
return node[1]
end
# is_side_effect_free_node(node) -> mirrors ir/lowering.rs's own
# is_side_effect_free_expr exactly, on this file's list-shaped AST instead
# of the native Expr enum: true for a shape that structurally CANNOT have a
# side effect of its own (a literal, arithmetic/comparison over such, a
# bare variable read, a list literal of such) -- everything that dispatches
# through a host/function call, a message send, or a closure creation is
# false, since those already speak for themselves via whatever they do; see
# lower_stmt's "Expr" arm below for why this distinction matters (auto-
# printing a Call's own incidental return value, e.g. print(x)'s Unit,
# would be noise, not the intended feature).
make a function called is_side_effect_free_node takes node returns r
let ty = node[0]
if (ty == "Num") or (ty == "Str") or (ty == "Bool") or (ty == "Var") then
return true
end
if ty == "Un" then
return is_side_effect_free_node(node[2])
end
if ty == "Bin" then
return is_side_effect_free_node(node[2]) and is_side_effect_free_node(node[3])
end
if ty == "List" then
let items = node[1]
let i = 0
let all_free = true
while i < items.length do
if is_side_effect_free_node(items[i]) == false then
let all_free = false
end
let i = i + 1
end
return all_free
end
return false
end
# wants_value: mirrors ir/lowering.rs's own `wants_value` thread exactly --
# true when this statement is in TAIL POSITION of a block/function that
# itself wants a value (the "blocks and functions return their last value"
# feature), meaning its own computed value should be left on the stack for
# the caller instead of discarded/auto-printed. See lower_block below for
# how a statement LIST decides which single statement (if any) gets
# wants_value=true.
# =============================================================================
# match/case (issue #44). Load-bearing design decision: `match` is PURE
# SYNTACTIC SUGAR -- compile_pattern below only ever produces ordinary AST
# nodes (Bin/Call/Index/Member/Var/Str/Num) that lower_expr/lower_stmt
# already know how to lower, and lower_match itself only ever emits the
# same Const/Load/Store/Jump/JumpIfFalse instructions an if/elif chain
# already uses (see lower_stmt's existing "If" arm just below, which this
# mirrors exactly). No new Instr variant -- see the plan's own design
# rationale (issue #44 follow-up) for why: every execution path
# (interp.patlang, codegen.patlang/codegen_x64.patlang) gets `match` for
# free since they never see anything but instructions they already handle.
#
# compile_pattern(pattern, scrutineeNode) -> [testExprNode, [letStmtNodes]]
# testExprNode: an ordinary boolean-valued AST expr node.
# letStmtNodes: ["Let", name, valueNode, false, false] nodes to run BEFORE
# the arm body, once the test has passed (a PBind's binding,
# or nested bindings from a PList's sub-patterns).
# =============================================================================
make a function called compile_pattern takes pattern, scrutineeNode returns r
let tag = pattern[0]
if tag == "PWild" then
return [["Bool", "true"], []]
else
if tag == "PBind" then
return [["Bool", "true"], [["Let", pattern[1], scrutineeNode, false, false]]]
else
if tag == "PLit" then
return [["Bin", "==", scrutineeNode, pattern[1]], []]
else
if tag == "PCmp" then
return [["Bin", pattern[1], scrutineeNode, pattern[2]], []]
else
if tag == "PGlob" then
return [["Call", "glob_match", [scrutineeNode, ["Str", pattern[1]]]], []]
else
if tag == "PList" then
return compile_list_pattern(pattern[1], scrutineeNode)
else
# Unreachable for a successfully-parsed pattern; a stray
# ["Err", ...] pattern (already surfaced separately by
# parser.patlang's pf_walk_pattern) falls back to an
# always-false test rather than crashing the lowerer.
return [["Bool", "false"], []]
end
end
end
end
end
end
end
make a function called compile_list_pattern takes subs, scrutineeNode returns r
let n = to_num(list_len(subs))
let test = ["Bin", "and",
["Bin", "==", ["Call", "type_of", [scrutineeNode]], ["Str", "list"]],
["Bin", "==", ["Member", scrutineeNode, "length"], ["Num", "" + n]]]
let lets = []
let i = 0
while i < n do
let elemNode = ["Index", scrutineeNode, ["Num", "" + i]]
let sub = compile_pattern(subs[i], elemNode)
let test = ["Bin", "and", test, sub[0]]
let j = 0
while j < sub[1].length do
let lets = list_push(lets, sub[1][j])
let j = j + 1
end
let i = i + 1
end
return [test, lets]
end
# Mirrors lower_stmt's "If" arm exactly, generalized to N arms with a
# guaranteed-fail tail when no arm's pattern (+ optional guard) matches --
# a runtime error naming the failure (via contract_check, the same
# design-by-contract host primitive already used for other "this should be
# unreachable" cases in this file, e.g. immutable-reassignment above), NOT
# a silent no-op: PatLang has no static type system to check exhaustiveness
# against, so an unmatched scrutinee with no `_` arm is a genuine bug at
# runtime, not a case to swallow quietly (decided explicitly in the issue
# #44 follow-up plan, not left implicit).
make a function called lower_match takes node, code, fns, fname, locals, mutables, pending, lines, wants_value returns out
let scrutineeExpr = node[1]
let arms = node[2]
# Store the scrutinee ONCE into a fresh synthesized local -- same
# reasoning as this file's other __closure_N/__budgeted_N synthesized
# names: re-evaluating an expression with side effects once per arm
# (e.g. `match next_event() do ...`) would be a real correctness bug,
# not just wasted work.
let tmp = next_match_name()
let code = lower_expr(scrutineeExpr, code, fns, locals, pending, lines)
vec_push(code, ["Store", tmp])
vec_push(locals, tmp)
vec_push(mutables, false)
let scrutineeVar = ["Var", tmp]
let endJumps = vec_new()
let i = 0
let n = to_num(list_len(arms))
while i < n do
let arm = arms[i]
let pattern = arm[0]
let guard = arm[1]
let body = arm[2]
let comp = compile_pattern(pattern, scrutineeVar)
let testNode = comp[0]
let letStmts = comp[1]
let code = lower_expr(testNode, code, fns, locals, pending, lines)
let jifPattern = vec_len(code)
vec_push(code, ["JumpIfFalse", 0])
# Bindings the pattern introduced (e.g. PBind's `let v = scrutinee`)
# must be emitted BEFORE the guard is evaluated, not folded into one
# `pattern_test and guard` expression -- a `when` guard is documented
# to reference bindings its own pattern just introduced (e.g. `case n
# when n > 100`), and `n` doesn't exist as a real local until this
# Store actually runs. Evaluating the AND as a single expression would
# load an unbound `n` while still building the LEFT side's truth value,
# silently reading Unit/stale-load instead of the real scrutinee no
# matter what the guard says. Found via a real failing run (guarded(200)
# returning "small:200" instead of "big:200"), not by inspection.
let j = 0
while j < letStmts.length do
let code = lower_stmt(letStmts[j], code, fns, fname, locals, mutables, pending, lines, false)
let j = j + 1
end
let jifGuard = -1
if guard != false then
let code = lower_expr(guard, code, fns, locals, pending, lines)
let jifGuard = vec_len(code)
vec_push(code, ["JumpIfFalse", 0])
end
let code = lower_block(body, code, fns, fname, locals, mutables, pending, lines, wants_value)
let jmp = vec_len(code)
vec_push(code, ["Jump", 0])
vec_push(endJumps, jmp)
let next_arm_pc = vec_len(code)
vec_set(code, jifPattern, ["JumpIfFalse", next_arm_pc])
if jifGuard >= 0 then
vec_set(code, jifGuard, ["JumpIfFalse", next_arm_pc])
end
let i = i + 1
end
# No arm matched: fatal, mirrors PatLang's existing "no try/catch, host
# errors are fatal" philosophy (contract_check's ok=false arg always
# errors -- see its use for immutable reassignment above -- so nothing
# after this point ever actually runs; the Store/Const that follow it
# exist only to keep the instruction stream's stack shape consistent
# with the matched-arm paths for any backend that reasons about it
# statically, matching the same dead-but-shape-correct pattern the
# immutable-reassignment path above already relies on).
vec_push(code, ["Const", "str", fname])
vec_push(code, ["Const", "str", "assert"])
vec_push(code, ["Const", "str", "match: no case matched the scrutinee value"])
vec_push(code, ["Const", "bool", "false"])
vec_push(code, ["CallHost", "contract_check", 4])
vec_push(code, ["Store", "__discard"])
if wants_value then
vec_push(code, ["Const", "unit", ""])
end
let end_pc = vec_len(code)
let i = 0
while i < vec_len(endJumps) do
vec_set(code, vec_get(endJumps, i), ["Jump", end_pc])
let i = i + 1
end
return code
end
make a function called lower_stmt takes node, code, fns, fname, locals, mutables, pending, lines, wants_value returns out
let ty = node[0]
# Record this statement's source line BEFORE lowering it, at the pc it's
# about to start emitting at -- only Let/Expr/Return/Assert/MemberAssign
# carry one (see parser.patlang's parse_stmt wrapper); node[node.length-1]
# is that line regardless of the node's own arity, since it was always
# appended as the trailing element there.
if (ty == "Let") or (ty == "Expr") or (ty == "Return") or (ty == "Assert") or (ty == "MemberAssign") then
vec_push(lines, [vec_len(code), node[node.length - 1]])
end
if ty == "Let" then
let valNode = node[2]
if (valNode[0] == "Call") and (valNode[1] == "activate") and (to_num(valNode[2].length) == 1) then
let code = lower_activate_stmt(valNode[2][0], code, fns, fname, locals, mutables, pending, lines)
else
let code = lower_expr(node[2], code, fns, locals, pending, lines)
end
let isReassign = node[3]
let isMut = node[4]
if isReassign then
if vec_contains(locals, node[1]) then
let flag = find_mut_flag(locals, mutables, node[1])
if not flag then
# Statically-known immutable reassignment: emit a contract_check
# that always fails at this point (mirrors the native Stage 0
# lowerer's approach -- lower_program is infallible here too, so
# this is enforced as a guaranteed-fail assertion at the
# reassignment site rather than a hard compile error).
vec_push(code, ["Const", "str", fname])
vec_push(code, ["Const", "str", "assert"])
vec_push(code, ["Const", "str", "cannot assign twice to immutable variable `" + node[1] + "` (declare it `let mut " + node[1] + "` to allow reassignment)"])
vec_push(code, ["Const", "bool", "false"])
vec_push(code, ["CallHost", "contract_check", 4])
end
vec_push(code, ["Store", node[1]])
else
# Not seen in this scope (e.g. a captured/outer-scope name this flat
# tracker can't see) -- fall back to the old permissive behaviour.
vec_push(code, ["Store", node[1]])
vec_push(locals, node[1])
vec_push(mutables, true)
end
else
# `let` / `let mut`: always allowed to introduce or shadow.
vec_push(code, ["Store", node[1]])
vec_push(locals, node[1])
vec_push(mutables, isMut)
end
# Ruby-style assignment-as-expression (mirrors ir/lowering.rs's own
# Stmt::Let arm): Store already consumed the value off the stack, so
# when this `let` is in tail position of a block that wants a value,
# reload it -- the assigned value IS the block's value.
if wants_value then
vec_push(code, ["Load", node[1]])
end
return code
else
if ty == "MemberAssign" then
# obj.prop = value -- lowers to send(obj, "set", prop, value), the
# same host call the object system already uses for explicit
# send("obj","set","prop",value) calls (matches ir/lowering.rs's
# native Stmt::MemberAssign arm). Previously called a nonexistent
# host fn "set" (3 args) instead -- never caught because
# MemberAssign was itself dead code in the native parser.rs until
# a real parser bug (bare '=' always consumed as equality before
# this ever got a chance) was fixed; nothing had ever exercised
# this lowering path before.
let code = lower_expr(node[1], code, fns, locals, pending, lines)
vec_push(code, ["Const", "str", "set"])
vec_push(code, ["Const", "str", node[2]])
let code = lower_expr(node[3], code, fns, locals, pending, lines)
vec_push(code, ["CallHost", "send", 4])
if wants_value == false then
vec_push(code, ["Store", "__discard"])
end
return code
else
if ty == "Expr" then
let exprNode = node[1]
if (exprNode[0] == "Call") and (exprNode[1] == "activate") and (to_num(exprNode[2].length) == 1) then
let code = lower_activate_stmt(exprNode[2][0], code, fns, fname, locals, mutables, pending, lines)
else
let code = lower_expr(node[1], code, fns, locals, pending, lines)
end
# GitHub #31 fix (mirrors ir/lowering.rs's Stmt::ExprStmt): every
# function body ends with an explicit Return, so a bare expression
# statement's pushed value is never actually consumed UNLESS this
# statement is in tail position of a block/function that wants it
# (wants_value) -- the interpreter's heap Vec-based operand stack
# tolerates leaving it regardless, but the x64 backend's real
# rsp-based stack overflows once a large loop repeats an unassigned
# call statement enough times (traced via WinDbg to
# expand_includes_at_depth's ~30,000-line loop). When unconsumed:
# auto-print (the "echoed only if not consumed" feature) ONLY for a
# structurally side-effect-free shape -- a Call (e.g. `activate PLAN`
# itself, or an ordinary print(x)) already speaks for itself via its
# own side effect, so it's silently discarded instead, exactly as
# before this feature existed (mirrors ir/lowering.rs's own
# discard_or_use / is_side_effect_free_expr split).
if wants_value then
return code
end
if is_side_effect_free_node(exprNode) then
vec_push(code, ["CallHost", "print", 1])
return code
end
vec_push(code, ["Store", "__discard"])
return code
else
if ty == "Print" then
let code = lower_expr(node[1], code, fns, locals, pending, lines)
vec_push(code, ["CallHost", "print", 1])
if wants_value == false then
vec_push(code, ["Store", "__discard"])
end
return code
else
if ty == "Return" then
let code = lower_expr(node[1], code, fns, locals, pending, lines)
vec_push(code, ["Return"])
return code
else
if ty == "Match" then
return lower_match(node, code, fns, fname, locals, mutables, pending, lines, wants_value)
else
if ty == "If" then
let code = lower_expr(node[1], code, fns, locals, pending, lines)
let jif = vec_len(code)
vec_push(code, ["JumpIfFalse", 0])
let code = lower_block(node[2], code, fns, fname, locals, mutables, pending, lines, wants_value)
let jmp = vec_len(code)
vec_push(code, ["Jump", 0])
vec_set(code, jif, ["JumpIfFalse", vec_len(code)])
# node[3] (the else branch) is always a list, `[]` when the
# source had no `else` -- lower_block's own empty-list handling
# already pushes Unit when wants_value, so the "no else but the
# false path needs a value too" case falls out for free.
let code = lower_block(node[3], code, fns, fname, locals, mutables, pending, lines, wants_value)
vec_set(code, jmp, ["Jump", vec_len(code)])
return code
else
if ty == "While" then
if wants_value then
# A while loop's value is Unit on zero iterations, otherwise
# its last-run iteration's tail value -- mirrors
# ir/lowering.rs's While arm exactly: push the zero-
# iteration default up front, then each iteration that
# actually runs pops the previous placeholder/prior value
# before pushing its own, so exactly one "current result"
# slot exists on the stack the whole time.
vec_push(code, ["Const", "unit", ""])
let start = vec_len(code)
let code = lower_expr(node[1], code, fns, locals, pending, lines)
let jif = vec_len(code)
vec_push(code, ["JumpIfFalse", 0])
vec_push(code, ["Store", "__discard"])
let code = lower_block(node[2], code, fns, fname, locals, mutables, pending, lines, true)
if budgeted_depth_get() > 0 then
# REAL BUG FOUND AND FIXED (caught by a real stack
# overflow, not review -- see budgeted_run's own
# x64 verification notes): CallHost always pushes
# exactly one result, but budget_check's own return
# value is never used here, and nothing was popping
# it -- on this wants_value=true path specifically it
# also broke the "exactly one current-result slot"
# invariant this loop's own header comment describes,
# stacking a second, permanent value on top of it
# every single iteration. A tight loop with thousands
# of iterations per budgeted() timeslice turns that
# into megabytes of unreclaimed stack in minutes.
vec_push(code, ["CallHost", "budget_check", 0])
vec_push(code, ["Store", "__discard"])
end
vec_push(code, ["Jump", start])
vec_set(code, jif, ["JumpIfFalse", vec_len(code)])
return code
else
let start = vec_len(code)
let code = lower_expr(node[1], code, fns, locals, pending, lines)
let jif = vec_len(code)
vec_push(code, ["JumpIfFalse", 0])
let code = lower_block(node[2], code, fns, fname, locals, mutables, pending, lines, false)
# Lexically inside a budgeted(ms) { ... } block: check the time
# budget just before looping back (mirrors lowering.rs's
# equivalent Stage 0 instrumentation). Store "__discard"
# pops budget_check's own unused return value -- see the
# wants_value=true branch just above for the real bug
# this fixes (a real stack overflow, not a review catch).
if budgeted_depth_get() > 0 then
vec_push(code, ["CallHost", "budget_check", 0])
vec_push(code, ["Store", "__discard"])
end
vec_push(code, ["Jump", start])
vec_set(code, jif, ["JumpIfFalse", vec_len(code)])
return code
end
else
if ty == "Assert" then
# contract_check(func_name, kind, text, ok) — args pushed in
# order, ok (the evaluated condition) last
vec_push(code, ["Const", "str", fname])
vec_push(code, ["Const", "str", node[1]])
vec_push(code, ["Const", "str", ast_to_source_text(node[2])])
let code = lower_expr(node[2], code, fns, locals, pending, lines)
vec_push(code, ["CallHost", "contract_check", 4])
if wants_value == false then
vec_push(code, ["Store", "__discard"])
end
return code
else
if ty == "RuleDecl" then
# Sugar: lowers to exactly the Instr sequence a hand-written
# rule_add(head_pred, [head_args...], [[pred,[args...]], ...])
# call already produces -- mirrors ir/lowering.rs's own
# Stmt::RuleDecl arm. Args are compile-time string TOKENS
# (a bare rule-head `X` is a logic-variable name, not a
# local-variable reference to evaluate).
let head_pred = node[1]
let head_args = node[2]
let body = node[3]
vec_push(code, ["Const", "str", head_pred])
let i = 0
while i < head_args.length do
vec_push(code, ["Const", "str", rule_arg_str(head_args[i])])
let i = i + 1
end
vec_push(code, ["BuildList", head_args.length])
let j = 0
while j < body.length do
let bodygoal = body[j]
let pred = bodygoal[0]
let args = bodygoal[1]
vec_push(code, ["Const", "str", pred])
let k = 0
while k < args.length do
vec_push(code, ["Const", "str", rule_arg_str(args[k])])
let k = k + 1
end
vec_push(code, ["BuildList", args.length])
vec_push(code, ["BuildList", 2])
let j = j + 1
end
vec_push(code, ["BuildList", body.length])
vec_push(code, ["CallHost", "rule_add", 3])
if wants_value == false then
vec_push(code, ["Store", "__discard"])
end
return code
else
if ty == "GoalDecl" then
# Sugar: lowers to exactly the Instr sequence a hand-
# written goal_def(NAME, [[pred,[args...]], ...]) call
# already produces -- mirrors ir/lowering.rs's
# Stmt::GoalDecl arm and the RuleDecl arm just above
# (dep args are compile-time string tokens too, same
# convention as a rule body's goal args).
let gname = node[1]
let deps = node[2]
vec_push(code, ["Const", "str", gname])
let j = 0
while j < deps.length do
let depgoal = deps[j]
let pred = depgoal[0]
let args = depgoal[1]
vec_push(code, ["Const", "str", pred])
let k = 0
while k < args.length do
vec_push(code, ["Const", "str", rule_arg_str(args[k])])
let k = k + 1
end
vec_push(code, ["BuildList", args.length])
vec_push(code, ["BuildList", 2])
let j = j + 1
end
vec_push(code, ["BuildList", deps.length])
vec_push(code, ["CallHost", "goal_def", 2])
if wants_value == false then
vec_push(code, ["Store", "__discard"])
end
return code
else
if ty == "ClassDecl" then
# Slice 1+2+3 of the classes/traits/inheritance
# feature (see the "synchronous-questing-metcalfe"
# plan) -- mirrors ir/lowering.rs's Stmt::ClassDecl
# arm exactly. Field defaults are real expressions,
# evaluated via the normal lower_expr path. Methods
# (Slice 2) become genuine closures via
# lower_closure_literal, each with an implicit
# leading "self" param, same shape as `when` blocks.
# Trait names (Slice 3) are compile-time string
# tokens, same convention as RuleDecl/GoalDecl dep
# args -- a trait reference is a registry lookup
# key, not an expression to evaluate.
let cname = node[1]
let cparent = node[2]
let cfields = node[3]
let cmethods = node[4]
let ctraits = node[5]
vec_push(code, ["Const", "str", cname])
vec_push(code, ["Const", "str", cparent])
let j = 0
while j < cfields.length do
let fld = cfields[j]
vec_push(code, ["Const", "str", fld[0]])
let code = lower_expr(fld[1], code, fns, locals, pending, lines)
vec_push(code, ["BuildList", 2])
let j = j + 1
end
vec_push(code, ["BuildList", cfields.length])
let j = 0
while j < cmethods.length do
let mth = cmethods[j]
let mname = mth[0]
let mparams = mth[1]
let mbody = mth[2]
vec_push(code, ["Const", "str", mname])
let full_params = ["self"]
let k = 0
while k < mparams.length do
let full_params = list_push(full_params, mparams[k])
let k = k + 1
end
let code = lower_closure_literal(full_params, mbody, code, fns, locals, pending, lines)
vec_push(code, ["BuildList", 2])
let j = j + 1
end
vec_push(code, ["BuildList", cmethods.length])
let j = 0
while j < ctraits.length do
vec_push(code, ["Const", "str", ctraits[j]])
let j = j + 1
end
vec_push(code, ["BuildList", ctraits.length])
vec_push(code, ["CallHost", "class_def", 5])
if wants_value == false then
vec_push(code, ["Store", "__discard"])
end
return code
else
# unknown statement: no code -- but a tail position
# that wants a value still needs something on the
# stack (mirrors ir/lowering.rs's own catch-all arm).
if wants_value then
vec_push(code, ["Const", "unit", ""])
end
return code
end
end
end
end
end
end
end
end
end
end
end
end
end
# Mirrors ir/lowering.rs's own lower_stmt_list exactly: lowers a statement
# LIST as a unit. Every statement but the last gets wants_value=false; the
# last inherits the list's own wants_value, so a value genuinely flows out
# of the block only from its tail position. An empty list that wants a
# value produces Unit (the block's value when it has no statements at all).
make a function called lower_block takes stmts, code, fns, fname, locals, mutables, pending, lines, wants_value returns out
if stmts.length == 0 then
if wants_value then
vec_push(code, ["Const", "unit", ""])
end
return code
end
let last = stmts.length - 1
let i = 0
while i < stmts.length do
let stmt_wants_value = (i == last) and wants_value
let code = lower_stmt(stmts[i], code, fns, fname, locals, mutables, pending, lines, stmt_wants_value)
let i = i + 1
end
return code
end
# ---- GitHub #102/#50: automatic box-on-crossing for ambiguous call
# arguments, --x64 only ----
#
# The boxed-Float REPRESENTATION (rt_float_tag/rt_box_float/rt_unbox_
# float_bits, x64_runtime.patlang; classification + bitwise-op unboxing,
# codegen_x64.patlang) already exists and is verified correct -- what
# was still missing is deciding WHERE to box automatically, so #102's
# actual repro (a function called with a float argument at one call
# site and a plain int at another) fixes itself without a caller having
# to call rt_box_float by hand.
#
# APPROACH (syntactic, not full dataflow -- limits stated below):
# a whole-program pre-scan, run once before lowering, finds every
# (function name, argument position) pair that is called with a
# SYNTACTICALLY float-looking argument expression at at least one call
# site and a NOT-float-looking one at another -- i.e. genuinely
# ambiguous across call sites, exactly #102's shape. Lowering a Call's
# argument for such a pair, when THIS specific argument expression also
# looks float, wraps it in a `rt_box_float` call; every other argument
# (including every argument of every non-ambiguous function) lowers
# completely unchanged, so a function that's consistently called with
# one type pays nothing extra -- the whole point of the "typed homes,
# box only at genuinely ambiguous crossings" design from #50.
#
# NAMED LIMITATION: "looks float" is judged SYNTACTICALLY on the
# argument expression itself (a dotted literal, arithmetic combining
# one, or a call to a known float-returning builtin) -- it does NOT
# trace data flow through a variable. `f(1.0 + 61.0)` is caught;
# `let x = 1.0 + 61.0; f(x)` is NOT (x's own taint isn't traced back to
# this call site). Closing that gap needs the real interprocedural
# taint-propagation extension discussed on #50/#102 -- genuinely
# separate, larger follow-up work, not attempted here. This pass also
# only looks at DIRECT `Call` argument expressions inside Let/Return/If/
# While conditions and bodies/MemberAssign/bare-expression statements --
# it does not descend into closures or List/BuildList/Index/Member
# sub-expressions beyond one hop, another named, deliberate scope cut
# for a first working version.
#
# Gated entirely on x64_lower_target_is_x64() (set by patc1_main.patlang
# only when --x64 is the active target) -- lower_program's output for
# every OTHER backend is completely unchanged; this whole mechanism is
# a no-op there.
make a function called x64_lower_target_is_x64 returns is_x64
let v = get("__vars", "x64_lower_target_is_x64")
return v == true
end
make a function called x64_lower_set_target_is_x64 takes flag returns done
set_var("x64_lower_target_is_x64", flag)
return true
end
# Syntactic float classifier -- see this section's own header for scope.
# `tainted` is a list of local variable names PROVEN float-looking
# somewhere earlier in the SAME function's own statement order (see
# ast_collect_float_tainted_vars_in_stmts) -- this is real, if simple,
# single-forward-pass dataflow (not just per-expression syntax): closes
# the `let x = 1.0 + 61.0; f(x)` case the original syntax-only version
# missed, since x's own taint is now looked up via the Var case below
# instead of unconditionally returning false for any Var.
make a function called ast_expr_looks_float_ctx takes expr, tainted returns is_float
let ty = expr[0]
if ty == "Num" then
return contains_str(expr[1], ".")
elif ty == "Var" then
return x64_list_contains_str(tainted, expr[1])
elif ty == "Bin" then
let op = expr[1]
if (op == "+") or (op == "-") or (op == "*") or (op == "/") or (op == "%") then
return ast_expr_looks_float_ctx(expr[2], tainted) or ast_expr_looks_float_ctx(expr[3], tainted)
end
return false
elif ty == "Un" then
if expr[1] == "-" then
return ast_expr_looks_float_ctx(expr[2], tainted)
end
return false
elif ty == "Call" then
let fname = expr[1]
if (fname == "sqrt") or (fname == "sin") or (fname == "cos") or (fname == "tan") or (fname == "asin") or (fname == "acos") or (fname == "atan") or (fname == "atan2") or (fname == "log") or (fname == "exp") or (fname == "pow") then
return true
end
return false
else
return false
end
end
# No-context convenience wrapper (equivalent to ast_expr_looks_float_ctx
# with an empty tainted-variable set) -- kept for any caller that
# genuinely has no local-scope context to offer (there are none left in
# this file after wiring the rewrite pass below, but keeping the name
# stable avoids an unnecessary rename of a small utility).
make a function called ast_expr_looks_float takes expr returns is_float
return ast_expr_looks_float_ctx(expr, [])
end
# ast_list_replace_at(lst, idx, newval) -> a NEW list, identical to lst
# except index idx is replaced with newval. Used by the AST rewrite pass
# below to change exactly one field of a statement/expression node
# (e.g. a Let's value expression) without having to know or reconstruct
# that node's full arity -- several of these AST node shapes carry
# extra trailing fields (Let's mutability flag, Func's optional named-
# return hint) that would be silently dropped by rebuilding the node
# from a fixed guessed field list instead.
make a function called ast_list_replace_at takes lst, idx, newval returns lst2
let out = []
let i = 0
while i < lst.length do
if i == idx then
let out = list_push(out, newval)
else
let out = list_push(out, lst[i])
end
let i = i + 1
end
return out
end
# Single forward pass over stmts: a `let x = <expr>` whose expr looks
# float (judged against the tainted set as grown so far, so `let y = x`
# after `let x = 1.0` also propagates) adds x to the returned set.
# Recurses into If/While bodies in the SAME scope (an assignment inside
# a branch still taints the name for code after the branch, matching
# this being a real, if conservative, over-approximation -- consistent
# with this whole pass's "syntactic, single forward pass, not a fixed-
# point dataflow solver" scope, named in this section's own header).
make a function called ast_collect_float_tainted_vars_in_stmts takes stmts, tainted returns tainted2
let result = tainted
let i = 0
while i < stmts.length do
let s = stmts[i]
let ty = s[0]
if ty == "Let" then
if ast_expr_looks_float_ctx(s[2], result) then
if not x64_list_contains_str(result, s[1]) then
let result = list_push(result, s[1])
end
end
elif ty == "If" then
let result = ast_collect_float_tainted_vars_in_stmts(s[2], result)
let result = ast_collect_float_tainted_vars_in_stmts(s[3], result)
elif ty == "While" then
let result = ast_collect_float_tainted_vars_in_stmts(s[2], result)
end
let i = i + 1
end
return result
end
# ---- AST rewrite pass: actually inserts the rt_box_float wrapper ----
#
# Runs AFTER x64_compute_ambiguous_call_params (needs its result) and
# BEFORE the real lower_program body processes stmts -- a genuine
# AST-to-AST transformation, not a codegen-time decision, precisely so
# each function's own tainted-variable context can be computed once and
# threaded naturally through recursion (lower_expr itself has no
# "which function/scope am I in" parameter to hang this off of).
# GitHub #99 regression found and fixed while completing #25 Piece 1:
# the box-on-crossing rewrite must ONLY ever apply to an ordinary Call
# to a ordinary user-visible PatLang FUNCTION (fns membership -- same
# gate lower_expr's own Call-lowering already uses to decide Call vs
# CallHost) -- NOT to a bare CallHost-dispatched primitive name like
# ceil/round/trunc, which take a raw float and would silently receive a
# boxed pointer instead if wrapped. `floor` is exactly this trap: it IS
# in `fns` (x64_runtime.patlang defines a real PatLang function with
# that name, for its own Rational case), so ceil(2.4)/round(2.4)/etc
# calls elsewhere in the SAME file, none of which are in `fns`, were
# never meant to be touched -- but the first version of this pass
# didn't check `fns` membership at all before deciding "ambiguous", so
# EVERY rounding-function call in a file mixing int/float rounding
# calls got its argument wrongly boxed. Found immediately re-testing
# self_hosting/examples/x64_rounding_ops.patlang (every single check
# failed, not just floor's) while verifying #99's real fix.
make a function called x64_rewrite_expr_boxing takes expr, ambiguous, tainted, fns returns expr2
let ty = expr[0]
if ty == "Call" then
let callee = expr[1]
let args = expr[2]
let new_args = []
let i = 0
while i < args.length do
let rewritten = x64_rewrite_expr_boxing(args[i], ambiguous, tainted, fns)
if contains_str(fns, callee) then
if not x64_list_contains_str(x64_box_crossing_excluded_names(), callee) then
if x64_list_contains_str(ambiguous, callee + "#" + i) then
if ast_expr_looks_float_ctx(args[i], tainted) then
let rewritten = ["Call", "rt_box_float", [rewritten]]
end
end
end
end
let new_args = list_push(new_args, rewritten)
let i = i + 1
end
return ast_list_replace_at(expr, 2, new_args)
elif ty == "Bin" then
let e2 = ast_list_replace_at(expr, 2, x64_rewrite_expr_boxing(expr[2], ambiguous, tainted, fns))
return ast_list_replace_at(e2, 3, x64_rewrite_expr_boxing(expr[3], ambiguous, tainted, fns))
elif ty == "Un" then
return ast_list_replace_at(expr, 2, x64_rewrite_expr_boxing(expr[2], ambiguous, tainted, fns))
else
return expr
end
end
make a function called x64_rewrite_stmts_boxing takes stmts, ambiguous, tainted, fns returns stmts2
let out = []
let cur_tainted = tainted
let i = 0
while i < stmts.length do
let s = stmts[i]
let ty = s[0]
if ty == "Let" then
if ast_expr_looks_float_ctx(s[2], cur_tainted) then
if not x64_list_contains_str(cur_tainted, s[1]) then
let cur_tainted = list_push(cur_tainted, s[1])
end
end
let new_s = ast_list_replace_at(s, 2, x64_rewrite_expr_boxing(s[2], ambiguous, cur_tainted, fns))
let out = list_push(out, new_s)
elif ty == "Return" then
if s.length > 1 then
let out = list_push(out, ast_list_replace_at(s, 1, x64_rewrite_expr_boxing(s[1], ambiguous, cur_tainted, fns)))
else
let out = list_push(out, s)
end
elif ty == "If" then
let s2 = ast_list_replace_at(s, 1, x64_rewrite_expr_boxing(s[1], ambiguous, cur_tainted, fns))
let s3 = ast_list_replace_at(s2, 2, x64_rewrite_stmts_boxing(s[2], ambiguous, cur_tainted, fns))
let s4 = ast_list_replace_at(s3, 3, x64_rewrite_stmts_boxing(s[3], ambiguous, cur_tainted, fns))
let out = list_push(out, s4)
elif ty == "While" then
let s2 = ast_list_replace_at(s, 1, x64_rewrite_expr_boxing(s[1], ambiguous, cur_tainted, fns))
let s3 = ast_list_replace_at(s2, 2, x64_rewrite_stmts_boxing(s[2], ambiguous, cur_tainted, fns))
let out = list_push(out, s3)
elif ty == "MemberAssign" then
let out = list_push(out, ast_list_replace_at(s, 3, x64_rewrite_expr_boxing(s[3], ambiguous, cur_tainted, fns)))
elif ty == "Expr" then
let out = list_push(out, ast_list_replace_at(s, 1, x64_rewrite_expr_boxing(s[1], ambiguous, cur_tainted, fns)))
elif ty == "Func" then
let fn_tainted = ast_collect_float_tainted_vars_in_stmts(s[3], [])
let new_body = x64_rewrite_stmts_boxing(s[3], ambiguous, fn_tainted, fns)
let out = list_push(out, ast_list_replace_at(s, 3, new_body))
else
let out = list_push(out, s)
end
let i = i + 1
end
return out
end
# Walks expr (and, one hop deep, its own Call-argument/Bin/Un operand
# sub-expressions) pushing [callee, arg_index, is_float] into the
# mutable Vec `out` for every direct Call it finds. `out` MUST be a Vec
# (vec_new()), not an ordinary list -- list_push is functional
# (returns a new list, leaves the caller's variable unchanged), which
# cannot accumulate across these recursive calls; vec_push mutates the
# handle in place, exactly what a shared accumulator threaded through
# recursion needs.
# `tainted` threads the enclosing function's own float-tainted local
# variables (see ast_collect_float_tainted_vars_in_stmts) through this
# walk, so an observation for `f(x)` correctly reflects whether x was
# ever assigned a float-looking value earlier in the SAME function --
# not just x's own (context-free) syntax, which is never float for a
# bare Var node.
make a function called x64_collect_call_arg_floatness_in_expr takes expr, out, tainted, fns returns done
let ty = expr[0]
if ty == "Call" then
let callee = expr[1]
let args = expr[2]
if contains_str(fns, callee) then
if not x64_list_contains_str(x64_box_crossing_excluded_names(), callee) then
let i = 0
while i < args.length do
let is_f = ast_expr_looks_float_ctx(args[i], tainted)
vec_push(out, [callee, i, is_f])
let i = i + 1
end
end
end
let i = 0
while i < args.length do
x64_collect_call_arg_floatness_in_expr(args[i], out, tainted, fns)
let i = i + 1
end
elif ty == "Bin" then
x64_collect_call_arg_floatness_in_expr(expr[2], out, tainted, fns)
x64_collect_call_arg_floatness_in_expr(expr[3], out, tainted, fns)
elif ty == "Un" then
x64_collect_call_arg_floatness_in_expr(expr[2], out, tainted, fns)
end
end
make a function called x64_collect_call_arg_floatness_in_stmts takes stmts, out, tainted, fns returns done
let cur_tainted = tainted
let i = 0
while i < stmts.length do
let s = stmts[i]
let ty = s[0]
if ty == "Let" then
x64_collect_call_arg_floatness_in_expr(s[2], out, cur_tainted, fns)
if ast_expr_looks_float_ctx(s[2], cur_tainted) then
if not x64_list_contains_str(cur_tainted, s[1]) then
let cur_tainted = list_push(cur_tainted, s[1])
end
end
elif ty == "Return" then
if s.length > 1 then
x64_collect_call_arg_floatness_in_expr(s[1], out, cur_tainted, fns)
end
elif ty == "If" then
x64_collect_call_arg_floatness_in_expr(s[1], out, cur_tainted, fns)
x64_collect_call_arg_floatness_in_stmts(s[2], out, cur_tainted, fns)
x64_collect_call_arg_floatness_in_stmts(s[3], out, cur_tainted, fns)
elif ty == "While" then
x64_collect_call_arg_floatness_in_expr(s[1], out, cur_tainted, fns)
x64_collect_call_arg_floatness_in_stmts(s[2], out, cur_tainted, fns)
elif ty == "MemberAssign" then
x64_collect_call_arg_floatness_in_expr(s[1], out, cur_tainted, fns)
x64_collect_call_arg_floatness_in_expr(s[3], out, cur_tainted, fns)
elif ty == "Func" then
x64_collect_call_arg_floatness_in_stmts(s[3], out, [], fns)
elif ty == "Expr" then
# A bare expression used as a statement (e.g. `bandtest(1)` on its
# own line) parses as ["Expr", <the real expression>, <line>], not
# the bare expression node directly -- found via a real repro that
# produced zero observations until this was dumped and inspected.
x64_collect_call_arg_floatness_in_expr(s[1], out, cur_tainted, fns)
else
x64_collect_call_arg_floatness_in_expr(s, out, cur_tainted, fns)
end
let i = i + 1
end
end
# GitHub #99 (found completing #25 Piece 1): a function that is ALREADY
# internally adaptive over int/float via the same runtime round-trip
# fixnum test x64_roundlike_asm/x64_abs_asm use (an int argument passes
# through unchanged, only a genuine float takes the real float path)
# must never be box-on-crossing eligible: boxing wraps its argument in
# a heap pointer, which that round-trip test cannot distinguish from
# "a very large but genuine int" -- it takes the float path on the
# pointer's own bits, computing garbage. `floor` is the concrete case:
# it's an ordinary PatLang-level Call (x64_runtime.patlang's own
# Rational-handling wrapper, colliding with the C import name), called
# with both float and plain-int arguments across a real program is
# exactly the ambiguity pattern this pass looks for, and it was WRONGLY
# getting boxed before this exclusion -- confirmed directly: floor(2.4)
# alone (no int call site anywhere else) computed correctly, but adding
# a single floor(7) call elsewhere in the same file broke it.
make a function called x64_box_crossing_excluded_names returns names
return ["floor"]
end
make a function called x64_list_contains_str takes xs, target returns found
let i = 0
while i < xs.length do
if xs[i] == target then
return true
end
let i = i + 1
end
return false
end
# Returns a list of "funcname#argindex" string keys for every (function,
# argument position) pair observed with BOTH a float-looking AND a
# not-float-looking argument across all its call sites in this program.
make a function called x64_compute_ambiguous_call_params takes stmts, fns returns keys
let observations = vec_new()
x64_collect_call_arg_floatness_in_stmts(stmts, observations, [], fns)
let n = vec_len(observations)
let saw_float = []
let saw_nonfloat = []
let i = 0
while i < n do
let obs = vec_get(observations, i)
let key = obs[0] + "#" + obs[1]
if obs[2] then
if not x64_list_contains_str(saw_float, key) then
let saw_float = list_push(saw_float, key)
end
else
if not x64_list_contains_str(saw_nonfloat, key) then
let saw_nonfloat = list_push(saw_nonfloat, key)
end
end
let i = i + 1
end
let keys = []
let j = 0
while j < saw_float.length do
if x64_list_contains_str(saw_nonfloat, saw_float[j]) then
let keys = list_push(keys, saw_float[j])
end
let j = j + 1
end
return keys
end
make a function called collect_fns takes stmts returns fns
let fns = []
let i = 0
while i < stmts.length do
let s = stmts[i]
if s[0] == "Func" then
let fns = list_push(fns, s[1])
end
let i = i + 1
end
return fns
end
# body_assigns_name(stmts, name) -> mirrors rust-runtime/src/parser.rs's
# own body_assigns_name exactly, on this file's list-shaped statement AST:
# true if `name` is `let`-assigned anywhere in `stmts`, at any nesting
# depth inside if/while branches (a shallow top-level-only scan would miss
# the common `if cond then let r = X else let r = Y end` shape). Used below
# to decide whether the `returns NAME` hint's synthesized trailing return
# should fire -- see that call site's own comment for why unconditionally
# synthesizing it would silently defeat implicit-last-value-return for any
# function that never touches NAME at all.
make a function called body_assigns_name takes stmts, name returns r
let i = 0
while i < stmts.length do
let s = stmts[i]
if (s[0] == "Let") and (s[1] == name) then
return true
end
if s[0] == "If" then
if body_assigns_name(s[2], name) then
return true
end
if body_assigns_name(s[3], name) then
return true
end
end
if s[0] == "While" then
if body_assigns_name(s[2], name) then
return true
end
end
if s[0] == "Match" then
let arms = s[2]
let ai = 0
while ai < arms.length do
if body_assigns_name(arms[ai][2], name) then
return true
end
let ai = ai + 1
end
end
let i = i + 1
end
return false
end
# Index of the LAST non-"Func" top-level statement, or -1 if there are
# none -- mirrors ir/lowering.rs's own `top_level`/`last_idx` computation
# in lower_program_basic: main's own tail statement wants a value too (the
# whole program's own implicit last-value return, restoring the pre-
# 86e4fef patc1_main echo behaviour as a consequence of this design rather
# than a separate code path), so its position needs to be known before the
# main loop below decides each statement's own wants_value.
make a function called last_nonfunc_index takes stmts returns idx
let last = -1
let i = 0
while i < stmts.length do
if stmts[i][0] != "Func" then
let last = i
end
let i = i + 1
end
return last
end
make a function called lower_program takes ast returns ir
set_var("__closure_seq", 0)
let stmts = ast[1]
let fns = collect_fns(stmts)
if x64_lower_target_is_x64() then
let ambiguous = x64_compute_ambiguous_call_params(stmts, fns)
let stmts = x64_rewrite_stmts_boxing(stmts, ambiguous, [], fns)
set_var("x64_ambiguous_call_params", ambiguous)
else
set_var("x64_ambiguous_call_params", [])
end
let funcs = []
let events = []
let mainCode = vec_new()
let mainLocals = vec_new()
let mainMutables = vec_new()
let mainLines = vec_new()
let pending = vec_new()
let handlerCount = 0
let last_top_level = last_nonfunc_index(stmts)
let i = 0
while i < stmts.length do
let s = stmts[i]
if s[0] == "Func" then
let flocals = vec_new()
let fmutables = vec_new()
let k = 0
while k < s[2].length do
vec_push(flocals, s[2][k])
vec_push(fmutables, true)
let k = k + 1
end
# Named-return hint (GitHub issue #5): if this Func node carries a
# 5th element (the `returns NAME` hint, non-empty), append a
# synthesized `["Return", ["Var", NAME]]` to the AST-level body
# BEFORE lowering -- matching Stage 0's rust-runtime/src/parser.rs
# (which appends the same synthesized return at parse time, not
# lowering time), so both backends' synthesis point stays
# structurally analogous. Reached only via genuine fall-through:
# any earlier explicit `return` already exits before this point.
# Only synthesized when the body actually assigns the hint name
# somewhere (see body_assigns_name above) -- unconditionally
# appending this for EVERY `returns NAME` declaration would
# silently steal the tail position from a function relying on
# implicit-last-value-return instead, returning an unassigned NAME.
let fbody = s[3]
if s.length > 4 then
if s[4] != "" then
if body_assigns_name(s[3], s[4]) then
let fbody = list_push(s[3], ["Return", ["Var", s[4]]])
end
end
end
let flines = vec_new()
let code = lower_block(fbody, vec_new(), fns, s[1], flocals, fmutables, pending, flines, true)
vec_push(code, ["Return"])
let funcs = list_push(funcs, ["FuncIR", s[1], s[2], vec_to_list(code), vec_to_list(flines)])
else
if s[0] == "When" then
# `when EVENT { ... }` now lowers to a genuine closure (reusing
# lower_closure_literal exactly, just with fixed auto-bound
# params ["event_name","event_data"] instead of user-supplied
# ones), registered at RUNTIME via register_event_handler --
# mirroring native lowering.rs's lower_when. Previously
# synthesized as an ISOLATED standalone function (no access to
# anything outer-scope) registered in a compile-time EventIR
# list, exactly the same gotcha found and fixed natively: a
# handler could never see an enclosing `let` (worked around by
# re-declaring the same name fresh inside every handler body).
# `mainLocals` (the CURRENT accumulated set at this exact point
# in program order, since this "When" case is handled inline in
# the SAME single sequential pass as everything else, not a
# separate pre-pass) is what lower_closure_literal captures --
# so this genuinely sees whatever's been declared before it.
vec_push(mainCode, ["Const", "str", s[1]])
let mainCode = lower_closure_literal(["event_name", "event_data"], s[2], mainCode, fns, mainLocals, pending, mainLines)
vec_push(mainCode, ["CallHost", "register_event_handler", 2])
if i != last_top_level then
vec_push(mainCode, ["Store", "__discard"])
end
else
let mainCode = lower_stmt(s, mainCode, fns, "main", mainLocals, mainMutables, pending, mainLines, i == last_top_level)
end
end
let i = i + 1
end
if last_top_level == -1 then
vec_push(mainCode, ["Const", "unit", ""])
end
vec_push(mainCode, ["Return"])
let funcs = list_push(funcs, ["FuncIR", "main", [], vec_to_list(mainCode), vec_to_list(mainLines)])
let pendingList = vec_to_list(pending)
let i = 0
while i < pendingList.length do
let funcs = list_push(funcs, pendingList[i])
let i = i + 1
end
return ["ProgramIR", "main", funcs, events]
end
make a function called iri_opt takes opts, key, fallback returns value
let i = 0
let n = list_len(opts)
while i < n do
let kv = opts[i]
if kv[0] == key then return kv[1] end
let i = i + 1
end
return fallback
end
make a function called iri_has_cap takes name returns found
let caps = host_caps()
let i = 0
while i < list_len(caps) do
if caps[i] == name then return true end
let i = i + 1
end
return false
end
# Source -> ["ok", ir] or ["err", message].
make a function called interp_compile takes src returns result
let ast = parse_program(tokenize(src))
let e = parse_find_error(ast[1])
if e[0] != "" then
return ["err", "parse error (line " + e[1] + "): " + e[0]]
end
return ["ok", lower_program(ast)]
end
make a function called interp_pat_path takes opts returns path
let p = iri_opt(opts, "pat_path", "")
if p != "" then return p end
let e = "" + getenv("PATLANG_PAT")
if e != "" then return e end
return "rust-runtime/target/release/pat.exe"
end
# "world" or "process" for these options on this runtime, "" if neither can run.
make a function called interp_backend takes opts returns name
let want = iri_opt(opts, "backend", "auto")
let has_world = iri_has_cap("world_swap")
let has_proc = iri_has_cap("subprocess") and file_exists(interp_pat_path(opts))
if want == "world" then
if has_world then return "world" end
return ""
end
if want == "process" then
if has_proc then return "process" end
return ""
end
let needs_proc = iri_opt(opts, "timeout_ms", 0) > 0 or iri_opt(opts, "stdin", "") != ""
if needs_proc and has_proc then return "process" end
if has_world then return "world" end
if has_proc then return "process" end
return ""
end
make a function called iri_failed takes message returns result
return ["", message, false, []]
end
# Does this pat.exe understand --quiet (drops the echo of the final value)?
# Probed once per pat_path and cached in a var.
make a function called iri_quiet_supported takes pat returns yes
let key = "__interp_run_quiet_" + pat
let cached = "" + get("__vars", key)
if cached != "" then return cached == "1" end
let probe = "self_hosting/build/_interp_run_probe.patlang"
write_file(probe, "print(\"q\")" + chr(10))
let r = exec_capture_io(pat, ["--quiet", "--ir-run", probe], "")
remove_file(probe)
let quiet = r[0] == "q" + chr(10)
if quiet then set_var(key, "1") else set_var(key, "0") end
return quiet
end
make a function called iri_tmp_path takes opts returns path
let n = to_num("0" + get("__vars", "__interp_run_seq")) + 1
set_var("__interp_run_seq", "" + n)
let dir = iri_opt(opts, "tmp_dir", "self_hosting/build")
return dir + "/_interp_run_" + now_ms() + "_" + n + ".patlang"
end
make a function called iri_run_process takes src, opts returns result
let pat = interp_pat_path(opts)
let tmp = iri_tmp_path(opts)
write_file(tmp, src)
let quiet = iri_quiet_supported(pat)
let args = ["--ir-run", tmp]
if quiet then let args = ["--quiet", "--ir-run", tmp] end
let timeout = iri_opt(opts, "timeout_ms", 0)
let r = exec_capture_io(pat, args, iri_opt(opts, "stdin", ""), timeout)
remove_file(tmp)
let out = r[0]
if quiet == false and r[2] then
# An unpatched pat.exe echoes the program's final value (Unit -> an empty
# line) after its own output.
if out.length > 0 then let out = substr(out, 0, out.length - 1) end
end
return [out, r[1], r[2], []]
end
# interp_run_full(src, opts) -> [stdout, stderr, ok, vfs_out]
make a function called interp_run_full takes src, opts returns result
let backend = interp_backend(opts)
if backend == "" then
return iri_failed("interp_run: no backend available for these options (host_caps: " + list_len(host_caps()) + " entries)")
end
let base_dir = iri_opt(opts, "base_dir", "")
if base_dir != "" then let src = expand_includes(src, base_dir) end
if backend == "process" then
return iri_run_process(src, opts)
end
if iri_opt(opts, "stdin", "") != "" then
return iri_failed("interp_run: stdin needs the process backend")
end
let c = interp_compile(src)
if c[0] == "err" then return iri_failed(c[1]) end
return interp_run_ir(c[1], opts)
end
# interp_run(src, opts) -> [stdout, stderr, ok]
make a function called interp_run takes src, opts returns result
let r = interp_run_full(src, opts)
return [r[0], r[1], r[2]]
end
# World backend on an already-lowered IR shape (skips compile: run the same
# program many times, or build IR directly).
make a function called interp_run_ir takes ir, opts returns result
if iri_has_cap("world_swap") == false then
return iri_failed("interp_run_ir: this runtime has no world_swap")
end
return world_run(ir, opts)
end
# Runs src on both backends and compares stdout, ok and the first line of
# stderr. Returns [same, detail]; only meaningful for programs that terminate.
make a function called interp_run_check takes src, opts returns result
let a = interp_run_full(src, list_push(opts, ["backend", "world"]))
let b = interp_run_full(src, list_push(opts, ["backend", "process"]))
if a[0] != b[0] then return [false, "stdout differs: world=[" + a[0] + "] process=[" + b[0] + "]"] end
if a[2] != b[2] then return [false, "ok differs: world=" + a[2] + " process=" + b[2]] end
let ea = iri_first_line(a[1])
let eb = iri_first_line(b[1])
if ea != eb then return [false, "stderr differs: world=[" + ea + "] process=[" + eb + "]"] end
return [true, ""]
end
make a function called iri_first_line takes s returns line
let h = str_intern(s)
let n = sc_len(h)
let i = 0
while i < n do
if sc_code(h, i) == 10 then return substr(s, 0, i) end
let i = i + 1
end
return s
end
# ---- grounding (same ^[A-Z]-is-a-logic-variable gotcha as synthesis.patlang) ----
make a function called synth2_ground takes term returns grounded
return "v:" + term
end
# ---- hypothesis construction ----
# One example query: [predicate, arg, expect_provable (true/false)].
make a function called synth2_query takes pred, arg, expect returns q
return [pred, arg, expect]
end
make a function called synth2_clauses_for takes target_pred, base_pred, chain_pred returns src
let sb = sb_new()
if base_pred != "" then
sb_push(sb, "rule " + target_pred + "(X) :- " + base_pred + "(X).\n")
end
if chain_pred != "" then
sb_push(sb, "rule " + target_pred + "(X) :- " + chain_pred + "(X, Y), " + target_pred + "(Y).\n")
end
return sb_str(sb)
end
# Build a full standalone PatLang script: background facts + candidate
# clauses + one print per query, "1" if solve()'s result matches the
# query's expectation (provable == expect), "0" otherwise.
make a function called synth2_build_test_script takes background_src, clause_src, queries returns src
let sb = sb_new()
sb_push(sb, background_src)
sb_push(sb, "\n")
sb_push(sb, clause_src)
sb_push(sb, "\n")
let i = 0
let n = to_num(list_len(queries))
while i < n do
let q = queries[i]
let pred = q[0]
let arg = q[1]
let expect = q[2]
sb_push(sb, "let sols_" + to_str_i(i) + " = solve(\"" + pred + "\", [\"" + synth2_ground(arg) + "\"])\n")
sb_push(sb, "let provable_" + to_str_i(i) + " = to_num(list_len(sols_" + to_str_i(i) + ")) > 0\n")
if expect then
sb_push(sb, "print(provable_" + to_str_i(i) + ")\n")
else
sb_push(sb, "print(not provable_" + to_str_i(i) + ")\n")
end
let i = i + 1
end
return sb_str(sb)
end
# No int-to-string host function exists; build small labels by repeating a
# marker character instead (only ever used for a handful of queries).
make a function called to_str_i takes i returns s
let sb = sb_new()
let k = 0
while k <= i do
sb_push(sb, "x")
let k = k + 1
end
return sb_str(sb)
end
# Run one hypothesis (background + candidate clauses) against all queries
# in a fresh subprocess; true only if every query's actual provability
# matched its expectation.
# Every query's print() prints "true" only when its actual provability
# matched its expectation (see synth2_build_test_script), so a hypothesis
# is accepted iff no query printed "false" -- one substring check across
# the whole captured transcript, no per-line parsing needed.
make a function called synth2_test_hypothesis takes background_src, clause_src, queries returns ok
let script = synth2_build_test_script(background_src, clause_src, queries)
let r = interp_run(script, [])
if r[2] == false then return false end
return not contains_text(r[0], "false")
end
# Split `text` into trimmed non-empty lines. Same character-scan idiom as
# lib/test.patlang's run_feature, reused here to parse a hypothesis-test
# transcript line by line instead of only checking it in aggregate.
make a function called synth2_split_lines takes text returns lines
let h = str_intern(text)
let n = sc_len(h)
let i = 0
let line = sb_new()
let lines = []
while i <= n do
let c = sc_code(h, i)
if (c == 10) or (c == -1) then
let raw = sb_str(line)
let line = sb_new()
if raw != "" then
let lines = list_push(lines, raw)
end
let i = i + 1
else
if c == 13 then
let i = i + 1
else
sb_push(line, sc_char(h, i))
let i = i + 1
end
end
end
return lines
end
# Same subprocess isolation as synth2_test_hypothesis, but returns the
# PER-QUERY pass/fail transcript (aligned index-for-index with `queries`)
# instead of collapsing it to one aggregate boolean -- lets a caller
# pinpoint WHICH example(s) an unsatisfying hypothesis got wrong, not just
# that it failed.
make a function called synth2_diagnose_hypothesis takes background_src, clause_src, queries returns results
let script = synth2_build_test_script(background_src, clause_src, queries)
let r = interp_run(script, [])
return synth2_pad_results(synth2_lines_to_results(r[0]), list_len(queries))
end
make a function called synth2_lines_to_results takes out returns results
let lines = synth2_split_lines(out)
let results = []
let i = 0
let n = to_num(list_len(lines))
while i < n do
let results = list_push(results, lines[i] == "true")
let i = i + 1
end
return results
end
# A script that crashed part-way prints fewer lines than there are queries;
# the queries it never reached did not pass.
make a function called synth2_pad_results takes results, want returns padded
let padded = results
while list_len(padded) < want do
let padded = list_push(padded, false)
end
return padded
end
# ---- metarule search ----
# Try, in order: each base predicate alone, then each (base, chain)
# predicate pair together (a chain clause alone can never terminate --
# recursion needs a base case, so that combination isn't in the search
# space). Returns [base_pred, chain_pred] of the first accepted
# hypothesis (chain_pred is "" for a base-only accept), or [] if nothing
# in the search space fits.
make a function called synth2_induce_chain takes target_pred, base_preds, chain_preds, background_src, queries returns winner
let bi = 0
let bn = to_num(list_len(base_preds))
while bi < bn do
let base_pred = base_preds[bi]
let base_only = synth2_clauses_for(target_pred, base_pred, "")
if synth2_test_hypothesis(background_src, base_only, queries) then
return [base_pred, ""]
end
let bi = bi + 1
end
let ci = 0
let cn = to_num(list_len(chain_preds))
while ci < cn do
let bi2 = 0
while bi2 < bn do
let base_pred = base_preds[bi2]
let chain_pred = chain_preds[ci]
let both = synth2_clauses_for(target_pred, base_pred, chain_pred)
if synth2_test_hypothesis(background_src, both, queries) then
return [base_pred, chain_pred]
end
let bi2 = bi2 + 1
end
let ci = ci + 1
end
return []
end
# All length-k sequences over candidates, with repetition (cartesian
# product candidates^k), in candidate-list order -- so a distractor
# predicate placed earlier in `candidates` gets tried (and rejected)
# before a correct predicate placed later, genuinely exercising
# discrimination rather than just never reaching the distractor.
make a function called synth4_sequences takes candidates, k returns seqs
if k == 0 then
return [[]]
end
let shorter = synth4_sequences(candidates, k - 1)
let result = []
let ci = 0
let cn = to_num(list_len(candidates))
while ci < cn do
let cand = candidates[ci]
let si = 0
let sn = to_num(list_len(shorter))
while si < sn do
let result = list_push(result, list_push(shorter[si], cand))
let si = si + 1
end
let ci = ci + 1
end
return result
end
# Build `rule target(X) :- R1(X, Y1), R2(Y1, Y2), ..., Rk(Y(k-1), Yk).`
# for a sequence of binary predicate names (length k = the chain depth).
make a function called synth4_chain_clause takes target_pred, pred_seq returns src
let sb = sb_new()
sb_push(sb, "rule " + target_pred + "(X) :- ")
let i = 0
let n = to_num(list_len(pred_seq))
let prev_var = "X"
while i < n do
if i > 0 then
sb_push(sb, ", ")
end
let next_var = "Y" + synth4_var_suffix(i)
sb_push(sb, pred_seq[i] + "(" + prev_var + ", " + next_var + ")")
let prev_var = next_var
let i = i + 1
end
sb_push(sb, ".\n")
return sb_str(sb)
end
# No int-to-string host function exists (see synthesis_recursive.
# patlang's to_str_i); build small distinct suffixes the same way.
make a function called synth4_var_suffix takes i returns s
let sb = sb_new()
let k = 0
while k <= i do
sb_push(sb, "x")
let k = k + 1
end
return sb_str(sb)
end
# Search depth 1..max_depth, and within each depth every predicate
# sequence in synth4_sequences order, for the first chain that proves
# every positive example and fails every negative one. Returns the
# winning predicate sequence, or [] if nothing in the search space (up to
# max_depth hops) fits.
make a function called synth4_induce_chain takes target_pred, candidate_preds, max_depth, background_src, queries returns winner
let depth = 1
while depth <= max_depth do
let seqs = synth4_sequences(candidate_preds, depth)
let i = 0
let n = to_num(list_len(seqs))
while i < n do
let seq = seqs[i]
let clause_src = synth4_chain_clause(target_pred, seq)
if synth2_test_hypothesis(background_src, clause_src, queries) then
return seq
end
let i = i + 1
end
let depth = depth + 1
end
return []
end
# ---- background facts as data ----
make a function called synth5_fact_unary takes pred, arg returns f
return [pred, arg]
end
make a function called synth5_fact_binary takes pred, arg1, arg2 returns f
return [pred, arg1, arg2]
end
make a function called synth5_facts_to_src takes facts returns src
let sb = sb_new()
let i = 0
let n = to_num(list_len(facts))
while i < n do
let f = facts[i]
if to_num(list_len(f)) == 2 then
sb_push(sb, "rule " + f[0] + "(\"" + synth2_ground(f[1]) + "\").\n")
else
sb_push(sb, "rule " + f[0] + "(\"" + synth2_ground(f[1]) + "\", \"" + synth2_ground(f[2]) + "\").\n")
end
let i = i + 1
end
return sb_str(sb)
end
# ---- bottom-up witness search ----
# Milestone 6: was "first match only" (a real limitation for any fact
# graph with actual branching -- e.g. alice `parent` of both bob and
# carol would only ever witness whichever fact happened to be listed
# first, silently ignoring the sibling). synth5_first_binary_hop is kept
# below unchanged as a baseline/regression reference; synth5_bfs_chain
# (which used it) is likewise kept unchanged for the same reason. Actual
# induction now goes through synth5_all_binary_hops/
# synth5_bfs_all_chains, which explore every witnessed branch.
# The first binary fact's predicate+successor found with arg1 == cur, or
# ["", ""] if none.
make a function called synth5_first_binary_hop takes facts, cur returns hop
let i = 0
let n = to_num(list_len(facts))
while i < n do
let f = facts[i]
if to_num(list_len(f)) == 3 then
if f[1] == cur then
return [f[0], f[2]]
end
end
let i = i + 1
end
return ["", ""]
end
# Walk outward from `start` up to max_depth binary hops, recording the
# predicate name used at each hop. Returns [] if not even one hop is
# witnessed -- the "no evidence at all for this example" case that later
# becomes a gap-diagnosis entry.
make a function called synth5_bfs_chain takes start, facts, max_depth returns pred_seq
let seq = []
let cur = start
let depth = 0
let stuck = false
while (depth < max_depth) and (not stuck) do
let hop = synth5_first_binary_hop(facts, cur)
if hop[0] == "" then
let stuck = true
else
let seq = list_push(seq, hop[0])
let cur = hop[1]
let depth = depth + 1
end
end
return seq
end
# All binary facts' [predicate, successor] pairs with arg1 == cur, not
# just the first -- the branching-aware replacement for
# synth5_first_binary_hop.
make a function called synth5_all_binary_hops takes facts, cur returns hops
let hops = []
let i = 0
let n = to_num(list_len(facts))
while i < n do
let f = facts[i]
if to_num(list_len(f)) == 3 then
if f[1] == cur then
let hops = list_push(hops, [f[0], f[2]])
end
end
let i = i + 1
end
return hops
end
make a function called synth5_list_contains takes xs, needle returns found
let i = 0
let n = to_num(list_len(xs))
let found = false
while i < n do
if xs[i] == needle then
let found = true
end
let i = i + 1
end
return found
end
# Safety cap on how many distinct witnessed chains one recursion level
# keeps -- a highly-branching fact graph could otherwise blow up the
# chain-set combinatorially depth over depth.
make a function called synth5_cap_chains takes chains, cap returns out
let n = to_num(list_len(chains))
if n <= cap then
return chains
end
let out = []
let i = 0
while i < cap do
let out = list_push(out, chains[i])
let i = i + 1
end
return out
end
# Recursive helper: every witnessed continuation from `cur`, up to
# `remaining_depth` further hops, as a list of predicate-sequences. A
# dead end (no further hop) or hitting remaining_depth == 0 both
# contribute the empty continuation []; a node already in `visited`
# (this path's own history) is excluded, so a cyclic fact graph
# terminates instead of looping forever.
make a function called synth5_extend_chains takes cur, facts, remaining_depth, visited returns chains
if remaining_depth == 0 then
return [[]]
end
let all_hops = synth5_all_binary_hops(facts, cur)
let hops = []
let hi = 0
let hn = to_num(list_len(all_hops))
while hi < hn do
let hop = all_hops[hi]
if not synth5_list_contains(visited, hop[1]) then
let hops = list_push(hops, hop)
end
let hi = hi + 1
end
if to_num(list_len(hops)) == 0 then
return [[]]
end
let results = []
let sub_visited = list_push(visited, cur)
let i = 0
let n = to_num(list_len(hops))
while i < n do
let hop = hops[i]
let sub_chains = synth5_extend_chains(hop[1], facts, remaining_depth - 1, sub_visited)
let si = 0
let sn = to_num(list_len(sub_chains))
while si < sn do
let full = list_push([hop[0]], sub_chains[si])
let results = list_push(results, synth5_flatten_chain(full))
let si = si + 1
end
let i = i + 1
end
return synth5_cap_chains(results, 25)
end
# list_push([hop_pred], sub_chain) above nests sub_chain as a single
# element rather than splicing its predicate names in -- flatten that one
# level back into a plain predicate-sequence.
make a function called synth5_flatten_chain takes nested returns flat
let flat = [nested[0]]
let rest = nested[1]
let i = 0
let n = to_num(list_len(rest))
while i < n do
let flat = list_push(flat, rest[i])
let i = i + 1
end
return flat
end
# Every distinct maximal witnessed predicate-sequence from `start`, up to
# max_depth hops (branching-aware). Returns [] only if `start` has no
# outgoing binary fact at all -- still the "no evidence for this example"
# signal synth5_lgg_from_positive checks for.
make a function called synth5_bfs_all_chains takes start, facts, max_depth returns chains
let raw = synth5_extend_chains(start, facts, max_depth, [])
let chains = []
let i = 0
let n = to_num(list_len(raw))
while i < n do
if to_num(list_len(raw[i])) > 0 then
let chains = list_push(chains, raw[i])
end
let i = i + 1
end
return chains
end
make a function called synth5_seq_eq takes a, b returns eq
let na = to_num(list_len(a))
let nb = to_num(list_len(b))
if na != nb then
return false
end
let i = 0
let eq = true
while i < na do
if a[i] != b[i] then
let eq = false
end
let i = i + 1
end
return eq
end
make a function called synth5_seq_list_contains takes seqs, needle returns found
let i = 0
let n = to_num(list_len(seqs))
let found = false
while i < n do
if synth5_seq_eq(seqs[i], needle) then
let found = true
end
let i = i + 1
end
return found
end
# Longest common prefix of two predicate-sequences -- anti-unification on
# sequences: where they agree, keep; the first point they diverge is
# where generalization has to stop being any more specific.
make a function called synth5_common_prefix takes a, b returns prefix
let out = []
let i = 0
let na = to_num(list_len(a))
let nb = to_num(list_len(b))
let lim = na
if nb < na then
let lim = nb
end
let stop = false
while (i < lim) and (not stop) do
if a[i] == b[i] then
let out = list_push(out, a[i])
let i = i + 1
else
let stop = true
end
end
return out
end
# Witness every positive example (via the full branching chain SET, not
# one greedily-chosen chain), then anti-unify (common-prefix) across all
# of them. Since an example can now witness several distinct chains, the
# running generalization is itself a SET of surviving candidate
# sequences: each existing candidate is matched against every one of the
# next example's chains and updated to whichever shared prefix is
# longest, so an early example's arbitrary-order chain list can no longer
# lock in the wrong branch the way a single greedy walk could.
#
# Returns [generalized_seq, missing_values, no_common_structure]:
# missing_values -- positive examples with NO witnessed chain at all
# (unchanged from before: a real evidence gap).
# no_common_structure -- true if every positive example DID have
# evidence, but their witnessed chains never converged on any shared
# structure at all (the candidate set emptied out mid-merge) -- a
# genuinely different situation from a missing fact: the evidence
# that exists doesn't agree with itself.
make a function called synth5_lgg_from_positive takes pos_values, facts, max_depth returns result
let missing = []
let candidates = []
let have_candidates = false
let no_common = false
let i = 0
let n = to_num(list_len(pos_values))
while i < n do
let v = pos_values[i]
let chains = synth5_bfs_all_chains(v, facts, max_depth)
if to_num(list_len(chains)) == 0 then
let missing = list_push(missing, v)
else
if not have_candidates then
let candidates = chains
let have_candidates = true
else
let merged = []
let ci = 0
let cn = to_num(list_len(candidates))
while ci < cn do
let cand = candidates[ci]
let best = []
let hi = 0
let hn = to_num(list_len(chains))
while hi < hn do
let p = synth5_common_prefix(cand, chains[hi])
if to_num(list_len(p)) > to_num(list_len(best)) then
let best = p
end
let hi = hi + 1
end
if to_num(list_len(best)) > 0 then
if not synth5_seq_list_contains(merged, best) then
let merged = list_push(merged, best)
end
end
let ci = ci + 1
end
let candidates = merged
if to_num(list_len(candidates)) == 0 then
let no_common = true
end
end
end
let i = i + 1
end
let seq = []
if to_num(list_len(candidates)) > 0 then
let best_i = 0
let bi = 0
let bn = to_num(list_len(candidates))
while bi < bn do
if to_num(list_len(candidates[bi])) > to_num(list_len(candidates[best_i])) then
let best_i = bi
end
let bi = bi + 1
end
let seq = candidates[best_i]
end
return [seq, missing, no_common]
end
# ---- top-level induction with gap diagnosis ----
# queries: milestone 2's synth2_query list (mixes positive and negative
# examples, [pred, arg, expect]).
#
# Returns one of:
# ["ok", pred_seq]
# ["no_witness", missing_values] -- one or more positive examples had no
# witnessed chain at all; the given background facts simply don't say
# anything connecting that example's value to anything else. This is
# the case worth surfacing back to the user as a targeted question:
# "scenario X expects <target>, but no background relation mentions X
# -- is a fact about X missing from the BDD scenarios/background?"
# ["conflict", offending] -- a generalized hypothesis was found and
# tested, but it mismatched one or more examples; `offending` is a
# list of [pred, arg, expect, note] entries, one per mismatch, with a
# human-readable note distinguishing "positive example not covered"
# (the LGG over-generalized away something this example needed) from
# "negative example wrongly covered" (the background facts don't yet
# distinguish this example from the positives -- suggests a missing
# fact that would rule it out, or a predicate the search wasn't told
# about).
# ["no_common_structure", pos_values] -- every positive example had
# witnessed evidence (unlike no_witness), but their witnessed chains
# never converged on any shared structure at all -- the evidence
# itself disagrees, not merely absent. Milestone 6: only reachable
# now that a positive example's full branching chain SET is
# considered rather than one greedily-picked chain.
make a function called synth5_induce takes target_pred, pos_values, queries, facts, max_depth returns diagnosis
let lgg_result = synth5_lgg_from_positive(pos_values, facts, max_depth)
let seq = lgg_result[0]
let missing = lgg_result[1]
let no_common = lgg_result[2]
if to_num(list_len(missing)) > 0 then
return ["no_witness", missing]
end
if no_common then
return ["no_common_structure", pos_values]
end
if to_num(list_len(seq)) == 0 then
return ["no_witness", pos_values]
end
let background_src = synth5_facts_to_src(facts)
let clause_src = synth4_chain_clause(target_pred, seq)
let results = synth2_diagnose_hypothesis(background_src, clause_src, queries)
let offending = []
let i = 0
let n = to_num(list_len(queries))
while i < n do
if not results[i] then
let q = queries[i]
let note = "negative example wrongly covered by the induced rule -- background facts don't yet distinguish it from the positives"
if q[2] then
let note = "positive example not covered -- the generalized rule (from other examples' witnessed chains) dropped something this example needed"
end
let offending = list_push(offending, [q[0], q[1], q[2], note])
end
let i = i + 1
end
if to_num(list_len(offending)) == 0 then
return ["ok", seq]
end
return ["conflict", offending]
end
make a function called synth5_join_quoted takes values returns s
let sb = sb_new()
let i = 0
let n = to_num(list_len(values))
while i < n do
if i > 0 then
sb_push(sb, ", ")
end
sb_push(sb, "\"" + values[i] + "\"")
let i = i + 1
end
return sb_str(sb)
end
# ---- turning a diagnosis into an actual question, not just structured data ----
# synth5_induce's ["ok", ...] / ["no_witness", ...] / ["conflict", ...]
# tuple is machine-usable but isn't yet a question anyone could act on --
# this is the missing half of the user's original request: "if we get a
# 'no hypothesis' result, we should probably be able to deduce what type
# of information was missing... so we can ask the user about the gaps".
# Renders a diagnosis into one plain-English question per gap, addressed
# to whoever wrote the BDD scenarios/background facts, so a caller (an
# interactive tool, a report, a log line) has literal text to show
# instead of a status tag.
make a function called synth5_format_diagnosis takes target_pred, diagnosis returns questions
let status = diagnosis[0]
let questions = []
if status == "ok" then
return questions
end
if status == "no_witness" then
let missing = diagnosis[1]
let i = 0
let n = to_num(list_len(missing))
while i < n do
let v = missing[i]
let q = "The scenarios say " + target_pred + "(\"" + v + "\") should hold, but no background fact connects \"" + v + "\" to anything at all. Is a fact about \"" + v + "\" missing from the background knowledge, or should this scenario be removed?"
let questions = list_push(questions, q)
let i = i + 1
end
return questions
end
if status == "no_common_structure" then
let values = diagnosis[1]
let q = "The positive examples (" + synth5_join_quoted(values) + ") all have background evidence, but their witnessed relation chains never agree on a shared structure. Do they really all belong in the same category, or is a background fact linking them to each other the same way missing?"
let questions = list_push(questions, q)
return questions
end
# status == "conflict"
let offending = diagnosis[1]
let i = 0
let n = to_num(list_len(offending))
while i < n do
let entry = offending[i]
let arg = entry[1]
let expect = entry[2]
let q = ""
if expect then
let q = "The induced rule does not cover \"" + arg + "\", even though the scenarios say " + target_pred + "(\"" + arg + "\") should hold. Its witnessed relation chain differs from the other positive examples -- does \"" + arg + "\" really belong in this category, or is a background fact linking it the same way as the others missing?"
else
let q = "The induced rule also matches \"" + arg + "\", but the scenarios say " + target_pred + "(\"" + arg + "\") should NOT hold. Is there a background fact that distinguishes \"" + arg + "\" from the positive examples that's missing, or is an additional condition (a predicate not offered to the search) needed?"
end
let questions = list_push(questions, q)
let i = i + 1
end
return questions
end
# Convenience wrapper: induce, then format -- returns [diagnosis,
# questions] so a caller gets both the structured result and the
# human-readable text in one call.
make a function called synth5_induce_and_ask takes target_pred, pos_values, queries, facts, max_depth returns outcome
let diagnosis = synth5_induce(target_pred, pos_values, queries, facts, max_depth)
let questions = synth5_format_diagnosis(target_pred, diagnosis)
return [diagnosis, questions]
end
# =============================================================================
# schema_synthesis_bridge: connects schema_bdd.patlang to the inductive-
# logic-programming engine (synthesis_lgg.patlang) -- implementation plan
# Phase 3, item 2.
#
# synth5_induce derives a rule that's correct with respect to its own
# training examples; it has no way to know about an independent
# constraint a schema states (a business rule, a policy, an invariant
# from a completely different part of the system). This bridge asserts an
# already-induced rule into the real logic engine, enumerates every value
# it proves true via solve(), and checks each one against a schema
# operation -- catching a rule that's logically correct on its training
# data but still produces a value the schema forbids.
#
# rule_add/solve (rust-runtime/src/ir/hosts.rs) are the real backward-
# chaining logic engine already used by synth5_induce's own domain
# (parent/grandparent-style facts); synth5_induce's own `facts` parameter
# is plain data for ITS internal graph search, not automatically visible
# to solve() -- background facts and the induced rule must be separately
# asserted here before solve() can see them at all.
# =============================================================================
# =============================================================================
# schema_bdd: attach a Z-notation-style schema (declared state + a general
# invariant, plus named operations each with a require/ensure precondition/
# postcondition) to a BDD feature, and check a concrete Given/When/Then
# scenario as a WITNESS of that schema -- catching a scenario whose own
# Given already violates an operation's precondition, before any of the
# scenario's own Then assertions even run.
#
# Design decisions, and why (see the implementation plan this file was
# built from for the full investigation):
#
# - A schema's invariant/require/ensure are ORDINARY, separately-compiled
# PatLang functions returning bool, referenced by name string and called
# via apply() -- never the literal `require`/`ensure` keywords. Those
# keywords lower to contract_check (rust-runtime/src/ir/hosts.rs), which
# is FATAL on violation (confirmed directly, and independently documented
# in self_hosting/lib/primitive_registry.patlang's own header) -- unusable
# for a check that must report a diagnosis and keep running, the same
# reason primitive_registry.patlang's own try_ wrappers never use
# require/ensure for their real failure path either.
#
# - apply() has no argument-spread form (confirmed in
# rust-runtime/src/ir/interpreter.rs): a function written once, generic
# over any schema's own number of state variables, cannot pass one
# positional argument per variable. State and inputs are therefore always
# bundled as a single list argument: invariant_fn(state_list),
# require_fn(state_list, input_list), ensure_fn(before_list, after_list,
# input_list).
#
# - Binding extends the existing set_var/get("__vars", ...) convention
# every Gherkin step in this codebase already uses (there is no
# parametrized step-matching mechanism anywhere to extend instead --
# step() dispatch is exact-literal-text, zero-argument, confirmed across
# every existing _bdd_demo.patlang file). A scenario's own Given/Then step
# functions call schema_bind_state/schema_bind_input directly.
#
# HARD RULE: self_hosting/lib/test.patlang's run_feature never resets bound
# vars between scenarios (only t_skipping/t_pending_tags are per-scenario).
# Every scenario's Given/Then must explicitly rebind every declared state
# variable and input itself, every time -- including restating an unchanged
# variable in Then. Do not rely on a value surviving from a previous
# scenario, and do not assume an unmentioned variable defaults to its
# "before" value in Then -- an omitted rebind reads back as the not-yet-
# bound sentinel (see schema_state_values below), which will correctly
# fail the check rather than silently pass, but the failure will look like
# a real violation unless this rule is followed.
# =============================================================================
# =============================================================================
# pmap: a minimal Map<K,V> as an association list of [key, value] pairs with
# linear-scan lookup -- the exact same idiom already proven in
# self_hosting/lib/depgraph.patlang's depgraph_map_get/depgraph_map_append
# (a file->list-of-files map), generalised here to an arbitrary value type
# and given a full put/get/has/remove/keys surface. PatLang has no native
# Map/Dict type usable from self-hosted code at this scale.
#
# Functional / return-new-collection style throughout, for the same reason
# as pset.patlang: schema_bdd needs an independent before/after snapshot of
# state, which an in-place-mutating map wouldn't give for free.
#
# pmap_get returns [] (an empty list) for a missing key -- the same
# not-found sentinel depgraph_map_get already uses, not a distinct "Unit"
# literal (PatLang's dialect has no source-level Unit/nil literal to write
# directly). This is INHERENTLY AMBIGUOUS if a real stored value could
# itself be an empty list: pmap_has is the only call that actually
# distinguishes "absent" from "present but happens to look like the
# sentinel" -- never infer presence from pmap_get's return value alone.
# =============================================================================
make a function called pmap_new returns m
return []
end
make a function called pmap_has takes m, key returns found
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return true
end
let i = i + 1
end
return false
end
# See the module-level warning above: [] means "not found" here, which is
# indistinguishable from a genuinely stored empty-list value. Call
# pmap_has first whenever that distinction matters.
make a function called pmap_get takes m, key returns value
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return m[i][1]
end
let i = i + 1
end
return []
end
make a function called pmap_put takes m, key, value returns m2
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return list_set(m, i, [key, value])
end
let i = i + 1
end
return list_push(m, [key, value])
end
make a function called pmap_remove takes m, key returns m2
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] != key then
let out = list_push(out, m[i])
end
let i = i + 1
end
return out
end
make a function called pmap_keys takes m returns keys
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
let out = list_push(out, m[i][0])
let i = i + 1
end
return out
end
make a function called pmap_size takes m returns n
return to_num(list_len(m))
end
# Single, process-wide registry (like Step's own new("Step", text)
# convention -- schemas and operations are inherently global, not
# per-instance, so one lazily-created Dict is simpler than
# primitive_registry.patlang's counter-named multi-registry support, which
# this doesn't need).
make a function called schema_registry returns registry
let existing = get("__vars", "schema_bdd_registry_obj")
if existing then
return existing
end
let registry = new("Dict", "schema_bdd_registry")
set_var("schema_bdd_registry_obj", registry)
return registry
end
# state_var_names: list of strings naming the schema's declared state.
# invariant_fn: name of a function taking ONE argument (the state values,
# in the same order as state_var_names) and returning bool.
make a function called schema_define takes name, state_var_names, invariant_fn returns done
send(schema_registry(), "set", name + "__schema", [state_var_names, invariant_fn])
return true
end
make a function called schema_lookup takes name returns entry
return get(schema_registry(), name + "__schema")
end
# input_names: list of strings naming the operation's declared inputs.
# require_fn: name of a function taking (state_values, input_values),
# returning bool -- the operation's precondition.
# ensure_fn: name of a function taking (before_values, after_values,
# input_values), returning bool -- the operation's postcondition.
make a function called schema_operation takes schema_name, op_name, input_names, require_fn, ensure_fn returns done
send(schema_registry(), "set", schema_name + "::" + op_name + "__op", [schema_name, input_names, require_fn, ensure_fn])
return true
end
make a function called schema_lookup_operation takes schema_name, op_name returns entry
return get(schema_registry(), schema_name + "::" + op_name + "__op")
end
# ---- binding: a scenario's own step functions call these ----
make a function called schema_bind_state takes schema, var_name, phase, value returns done
set_var(schema + "__" + var_name + "__" + phase, value)
return true
end
make a function called schema_bind_input takes schema, op, param_name, value returns done
set_var(schema + "__" + op + "__in__" + param_name, value)
return true
end
# Not-yet-bound sentinel: get("__vars", ...) on an unset key -- same
# absence convention already relied on throughout this codebase (e.g.
# gherkin_contracts.patlang's `already_bound` check). Never distinguishes
# "never bound" from "bound to this same falsy value"; low risk for the
# LibraryLoans-shaped schemas this was designed against (string/list-typed
# state), a documented hard limit for anything boolean- or zero-valued.
make a function called schema_state_values takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__before"))
let i = i + 1
end
return out
end
make a function called schema_state_values_after takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__after"))
let i = i + 1
end
return out
end
make a function called schema_input_values takes schema, op, input_names returns values
let out = []
let i = 0
let n = to_num(list_len(input_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + op + "__in__" + input_names[i]))
let i = i + 1
end
return out
end
# ---- the witness check ----
#
# Four-stage check, each stage its own diagnosis tag, checked in the order
# a Z schema's own reasoning goes: is the state even valid to start from,
# does this operation's precondition actually hold for it, does the
# claimed resulting state stay valid, and does the operation's own
# postcondition connect before to after correctly. Returns [tag, payload],
# the same 2-element diagnosis shape as synth5_induce
# (self_hosting/lib/synthesis_lgg.patlang) -- deliberately, so
# schema_format_diagnosis below can mirror synth5_format_diagnosis's own
# real structure rather than inventing a new diagnosis style.
# The value-based core: takes state_before/inputs/state_after directly
# rather than reading them out of the __vars binding store. schema_check
# below is the BDD-scenario-shaped wrapper around this; this function is
# what any OTHER caller (a synthesis harness, a future GOAP or ILP hook --
# see the implementation plan's Phase 3) should call directly instead of
# faking a scenario's Given/When/Then bindings just to reach schema_check.
make a function called schema_check_values takes schema_name, op_name, state_before, inputs, state_after returns diagnosis
let schema_entry = schema_lookup(schema_name)
let invariant_fn = schema_entry[1]
let op_entry = schema_lookup_operation(schema_name, op_name)
let require_fn = op_entry[2]
let ensure_fn = op_entry[3]
if apply(invariant_fn, state_before) == false then
return ["invariant_violated_before", [schema_name, state_before]]
end
if apply(require_fn, state_before, inputs) == false then
return ["precondition_violated", [op_name, state_before, inputs]]
end
if apply(invariant_fn, state_after) == false then
return ["invariant_violated_after", [schema_name, state_after]]
end
if apply(ensure_fn, state_before, state_after, inputs) == false then
return ["postcondition_violated", [op_name, state_before, state_after, inputs]]
end
return ["ok", [op_name, state_before, state_after, inputs]]
end
make a function called schema_check takes schema_name, op_name returns diagnosis
let schema_entry = schema_lookup(schema_name)
let state_var_names = schema_entry[0]
let op_entry = schema_lookup_operation(schema_name, op_name)
let input_names = op_entry[1]
let state_before = schema_state_values(schema_name, state_var_names)
let inputs = schema_input_values(schema_name, op_name, input_names)
let state_after = schema_state_values_after(schema_name, state_var_names)
return schema_check_values(schema_name, op_name, state_before, inputs, state_after)
end
# Mirrors synth5_format_diagnosis's real structure (synthesis_lgg.patlang):
# switch on diagnosis[0], build a human-readable question per case. Returns
# a plain list of question strings (empty for "ok"), same shape as
# synth5_induce_and_ask's own return convention.
make a function called schema_format_diagnosis takes schema_name, op_name, diagnosis returns questions
let tag = diagnosis[0]
let payload = diagnosis[1]
if tag == "invariant_violated_before" then
return ["This scenario's Given already leaves " + schema_name + " in a state that violates its own invariant, before " + op_name + " is even checked. Is the Given wrong, or is the invariant too strict?"]
end
if tag == "precondition_violated" then
return ["Scenario claims " + op_name + " can run from this Given, but " + op_name + "'s own precondition returned false for these inputs. Is the Given wrong, or is " + op_name + "'s precondition too strict -- or is this scenario meant to test a rejection path, which needs its own operation schema rather than " + op_name + "'s?"]
end
if tag == "invariant_violated_after" then
return ["Applying " + op_name + " produces a state that violates " + schema_name + "'s invariant. Is the Then clause's claimed resulting state wrong, or does " + op_name + " need a stronger precondition to rule this case out?"]
end
if tag == "postcondition_violated" then
return ["The before/after state this scenario claims for " + op_name + " does not satisfy its own postcondition. Is the Then clause's claimed resulting state wrong, or is " + op_name + "'s postcondition wrong?"]
end
return []
end
# ---- synthesis integration (implementation plan Phase 3) ----
#
# Opt-in registration linking a SYNTHESIZED function's name to a schema
# operation it's meant to satisfy, plus a "harness" function name that
# knows how to actually exercise it: harness_fn takes (func_name) and
# returns [state_before, inputs, state_after] by calling the synthesized
# function itself (via apply) on some concrete test input and observing
# what it does. This is necessarily domain-specific -- there is no way to
# derive it generically -- so it must be supplied by whoever registers the
# hook, not synthesized here.
#
# Deliberately separate from schema_operation itself: a schema operation
# describes the CONTRACT; a synthesis hook additionally says "and here's
# how to actually run a candidate implementation against it," which only
# matters once there's a synthesized candidate to check, not for the
# scenario-witnessing use in schema_check above.
make a function called schema_register_synthesis_check takes func_name, schema_name, op_name, harness_fn returns done
send(schema_registry(), "set", func_name + "__synthesis_check", [schema_name, op_name, harness_fn])
return true
end
make a function called schema_lookup_synthesis_check takes func_name returns entry
return get(schema_registry(), func_name + "__synthesis_check")
end
# Runs a registered synthesis hook for func_name and returns its
# schema_check_values diagnosis directly. Callers with no hook registered
# for func_name should skip calling this entirely (schema_lookup_
# synthesis_check(func_name) is falsy) rather than call it and inspect the
# result -- there is no "no hook registered" diagnosis tag, since this
# function assumes a hook exists.
make a function called schema_run_synthesis_check takes func_name returns diagnosis
let hook = schema_lookup_synthesis_check(func_name)
let schema_name = hook[0]
let op_name = hook[1]
let harness_fn = hook[2]
let triple = apply(harness_fn, func_name)
return schema_check_values(schema_name, op_name, triple[0], triple[1], triple[2])
end
# facts: list of [pred, arg1] or [pred, arg1, arg2] tuples -- the same
# shape synth5_facts_to_src (synthesis_lgg.patlang) already renders as
# source text; asserted here directly via rule_add instead.
make a function called schema_assert_facts takes facts returns done
let i = 0
let n = to_num(list_len(facts))
while i < n do
let f = facts[i]
if to_num(list_len(f)) == 2 then
rule_add(f[0], [f[1]], [])
else
rule_add(f[0], [f[1], f[2]], [])
end
let i = i + 1
end
return true
end
# Mirrors synth4_chain_clause's own "rule target(X) :- R1(X,Y1),
# R2(Y1,Y2), ..." shape (synthesis_chain.patlang:52), asserted directly
# via rule_add instead of compiled from generated source text -- valid
# in-process for checking one already-chosen rule (no isolation concern,
# unlike synth2_test_hypothesis's fresh-subprocess-per-candidate search).
make a function called schema_assert_chain_rule takes target_pred, pred_seq returns done
let body = []
let i = 0
let n = to_num(list_len(pred_seq))
let prev_var = "X"
while i < n do
let next_var = "Y" + (i + 1)
let body = list_push(body, [pred_seq[i], [prev_var, next_var]])
let prev_var = next_var
let i = i + 1
end
rule_add(target_pred, ["X"], body)
return true
end
# Enumerates every value the already-asserted rule proves true for
# target_pred (via solve), maps each through mapper_fn -- a caller-
# supplied function taking one candidate value and returning
# [state_before, inputs, state_after], the same triple shape a synthesis
# harness returns in schema_bdd.patlang's schema_run_synthesis_check --
# and checks it against schema_name/op_name. Returns the FIRST non-"ok"
# diagnosis found, as [tag, [candidate_value, original_payload]], or
# ["ok", []] if every candidate the rule proves true also satisfies the
# schema.
make a function called schema_check_induced_rule takes schema_name, op_name, target_pred, mapper_fn returns diagnosis
let solutions = solve(target_pred, ["X"])
let i = 0
let n = to_num(list_len(solutions))
while i < n do
let candidate = solutions[i][0]
let triple = apply(mapper_fn, candidate)
let d = schema_check_values(schema_name, op_name, triple[0], triple[1], triple[2])
if d[0] != "ok" then
return [d[0], [candidate, d[1]]]
end
let i = i + 1
end
return ["ok", []]
end
# ---- GrandparentPolicy schema: state = [restricted], a pset of names
# certification is never allowed to name. No general invariant beyond
# "the state exists" -- the real constraint is entirely in Certify's own
# postcondition, deliberately independent of how a candidate was derived.
make a function called gp_inv_trivial takes state returns ok
return true
end
make a function called gp_require_true takes state, inputs returns ok
return true
end
make a function called gp_ensure_not_restricted takes before, after, inputs returns ok
let restricted = before[0]
return pset_contains(restricted, inputs[0]) == false
end
# mapper_fn for schema_check_induced_rule: state before/after both name
# "dave" as the one restricted name (the point is the policy, which never
# changes across the check, not any real state transition).
make a function called gp_mapper takes candidate returns triple
let restricted = pset_add(pset_new(), "dave")
return [[restricted], [candidate], [restricted]]
end
t_init()
schema_define("GrandparentPolicy", ["restricted"], "gp_inv_trivial")
schema_operation("GrandparentPolicy", "Certify", ["who"], "gp_require_true", "gp_ensure_not_restricted")
# ---- real induction, same domain/shape as synthesis_lgg_selftest.patlang's
# scenario A, with dave's line ending in "eve" instead of "frank" so both
# candidates are otherwise unremarkable to the induction engine itself.
let facts = [
synth5_fact_binary("parent", "alice", "bob"),
synth5_fact_binary("parent", "bob", "carol"),
synth5_fact_binary("parent", "dave", "erin"),
synth5_fact_binary("parent", "erin", "eve")
]
let queries = [
synth2_query("grandparent", "alice", true),
synth2_query("grandparent", "dave", true)
]
let induced = synth5_induce("grandparent", ["alice", "dave"], queries, facts, 2)
check("induction succeeds on the training examples", induced[0], "ok")
check("induced rule is the expected 2-hop parent/parent chain", induced[1], ["parent", "parent"])
# ---- bridge: assert the same facts + induced rule into the real logic
# engine, then check every value the rule proves true against the policy.
schema_assert_facts(facts)
schema_assert_chain_rule("grandparent", induced[1])
let bridge_diagnosis = schema_check_induced_rule("GrandparentPolicy", "Certify", "grandparent", "gp_mapper")
check("the induced rule is caught violating an unrelated schema policy", bridge_diagnosis[0], "postcondition_violated")
check("the flagged candidate is the restricted name, not the other one", bridge_diagnosis[1][0], "dave")
t_report()
(not run yet)
The Negative Case: Catching a Scenario Before It Runs Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Run the first example (one operation, one violation) and look at its second scenario: Given already states that “Dune” is borrowed by S. Okonkwo, then When has A. Diallo try to borrow it too. Checked against BorrowBook’s schema rather than run against any implementation, this fails before any code executes at all — its own Given already violates the operation’s precondition. That’s the mechanism’s whole point: schema_check returns a tagged diagnosis, ["precondition_violated", [op, state_before, inputs]], and schema_format_diagnosis turns it into a real, specific question rather than a bare pass/fail — mirroring the same tagged-diagnosis-plus-question pattern BDD as Specification’s induction engine already uses for a scenario a rule can’t satisfy:
> schema_check("LibraryLoans", "BorrowBook")
["precondition_violated", ["BorrowBook", [...], ["Dune", "A. Diallo"]]]
> schema_format_diagnosis("LibraryLoans", "BorrowBook", diagnosis)
["Scenario claims BorrowBook can run from this Given, but BorrowBook's own
precondition returned false for these inputs. Is the Given wrong, or is
BorrowBook's precondition too strict -- or is this scenario meant to test a
rejection path, which needs its own operation schema rather than
BorrowBook's?"]
That question is the actual payoff, not a rhetorical flourish. It distinguishes two mistakes that look identical from the outside — a scenario reporting an outcome the current implementation doesn’t produce — but need entirely different fixes: an operation modelled with the wrong precondition, versus a scenario silently testing a code path no schema was ever written to cover. A plain Given/When/Then runner can’t tell these apart; a schema, checked independently of any implementation, can.
A Fuller Worked Example: Two Operations and a Real Invariant Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The second dropdown option scales the same idea up: three pieces of interacting state (books, borrowedBy, and a per-member loanCount), two operations (BorrowBook, ReturnBook), and a genuine cross-cutting invariant — no member may hold more than three books at once, a rule that isn’t derivable from the book-tracking half of the state at all. Run it and its seven scenarios exercise four of the five outcomes schema_check can return (ok, precondition_violated, invariant_violated_before, and postcondition_violated; the fifth, invariant_violated_after, has no scenario here): an available title accepted, a title someone else has out rejected, a member at the loan limit rejected, a corrupted loan record caught by the invariant before the operation is even considered, a scenario whose own claimed effect doesn’t match what its postcondition demands, and a clean return accepted and a spurious one rejected. Full source: self_hosting/examples/library_loans_schema_demo.patlang.
Beyond Hand-Written Scenarios: Synthesis Integration
A schema doesn’t only check scenarios a person wrote by hand. The same schema_check_values core also plugs into two of PatLang’s existing program-synthesis mechanisms as a stronger correctness oracle — catching an artifact that’s internally consistent with its own training data or its own plan, but still wrong by an independent standard. Both are in the example list at the top of the page and run there, in the browser; the output each one produces is printed below its description.
Inductive Synthesis: A Logically Valid Rule, Still Policy-Violating Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
BDD as Specification covers PatLang’s inductive-logic-programming engine deriving grandparent(X) :- parent(X,Y), parent(Y,Z) from training facts and examples. That derivation is provably correct with respect to its own training data — but the engine has no way to know about a constraint from a completely different part of a system, e.g. a policy that names specific people. self_hosting/lib/schema_synthesis_bridge.patlang asserts an already-induced rule into PatLang’s logic engine, enumerates every value it proves true, and checks each one against a schema. To run it here, choose “Inductive synthesis” in the example list at the top of the page. The engine tests each candidate rule in a clean interpreter created inside the page: interp_run compiles the test script and hands it to world_run. That host call swaps every global store the interpreter keeps (facts, rules, objects, virtual files, string and list handles) for an empty set, runs the script with its output captured, and puts the caller’s state back afterwards. Native PatLang uses the same call instead of spawning pat.exe, and falls back to a subprocess only when a wall-clock timeout or stdin is asked for.
ok: induction succeeds on the training examples ok: induced rule is the expected 2-hop parent/parent chain ok: the induced rule is caught violating an unrelated schema policy ok: the flagged candidate is the restricted name, not the other one tests: 4 passed, 0 failed ALL TESTS PASSED
Two people, “alice” and “dave”, both satisfy the induced rule — the training data treats them identically. A GrandparentPolicy schema with an independent restricted-names list catches “dave” specifically, without touching the induction engine’s own logic at all. Full source: self_hosting/schema_synthesis_bridge_selftest.patlang.
GOAP Planning: Checking Real State, Not a Parsed Label Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
PatLang’s pre-existing GOAP contract system (goap_verify_contracts) can only check a contract against a string-parsed action-label binding, like extracting X=5 out of the text "scale(X=5)" — it has no way to express “and the book must not also still be at the origin branch,” because that needs the plan’s full resulting state, not one action’s own parameter. A new host function, plan_with_state, exposes that resulting state directly. It exists in three places: as a Rust interpreter host function, in the Rust-to-native codegen path (the one the in-page compiler is built from), and, canonically, as self-hosted PatLang in self_hosting/lib/x64_runtime.patlang, compiled and run through the patc1.exe --x64 production toolchain as well as the interpreter. Choose “GOAP: interlibrary transfer” in the example list at the top of the page to run it here.
ok: the planner finds the full three-hop route ok: step 1 packs the book for transit ok: step 2 ships it to the depot ok: step 3 ships it on to the destination branch ok: the real resulting state has the book at branch_b ok: the real resulting state no longer has it at branch_a ok: the real resulting state has no dangling in-transit record ok: the full transfer plan satisfies the transfer policy ok: a destination with no shipping route is reported as no_plan_found tests: 9 passed, 0 failed ALL TESTS PASSED
A rare book moves from one branch to another through a three-step GOAP plan (pack for transit, ship to the depot, ship on to the destination); the schema checks the plan’s resulting facts — the book present at the destination and absent from the origin, not inferred from the last action’s own label text. A destination with no shipping route is correctly reported as no_plan_found, distinct from a schema violation. Full source: self_hosting/examples/interlibrary_transfer_goap_demo.patlang.
The Other Direction: Schema Suggesting Scenarios Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Checking existing scenarios against a schema is the easier direction. The harder, more useful direction runs the other way: since a schema states a universally-quantified property rather than one instance, a schema-aware tool could enumerate concrete cases a scenario set doesn’t cover and propose them, rather than waiting for someone to think of the edge case by hand. This isn’t a new idea — it’s what property-based testing already does, generating concrete test cases from a stated property instead of a human enumerating them by hand3. This direction is now built, deterministically, with no LLM in the loop: given an operation, the generator searches the schema’s own reachable states and declared input domains — the same search the explorer already runs — for one state and input where the operation is enabled, and one per require clause where every other clause holds and that one alone fails. A clause with no such state, typically because it is implied by another or is tautological within its own declared domain, is reported rather than given an invented scenario.state the rule, generate the cases: cf. property-based testing
Every word the generator prints is either the schema author’s own text (a fact, then, or reject template, an operation or clause name) or a value taken from the state that was found and independently checked; nothing fills a gap in the schema’s own vocabulary. Run on the vending machine above, for BuyVM:
Feature: BuyVM (VendingMachine) Someone wants to buy a snack with money Scenario: BuyVM succeeds When cookie is in stock And Bowen selects cookie Then BuyVM succeeds Scenario: BuyVM is rejected: map_get_or(inventory, item, 0) > 0 When Bowen selects chips Then the vending machine reports map_get_or(inventory, item, 0) > 0 Scenario: BuyVM is rejected: currentMoneyValue >= map_get(price, item) When cookie is in stock And Bowen selects cookie Then the vending machine reports currentMoneyValue >= map_get(price, item)
The third require clause, item in snacks, is reported separately rather than given a scenario: the operation’s own domain for item is snacks, so every value the search can try already satisfies it, and inventing a value outside that domain to violate it would not be a value the schema was ever checked against. Feeding this generated feature back through the same analysis that checks a hand-written one reports it consistent — the generator’s own witnesses hold up under the same check a person’s scenario would face, which is the actual test of whether generating them was worth doing.
From Examples to Every Reachable State Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The mechanism above checks the states a scenario supplies. A second layer, written after reading Liu’s thesis, declares the schema as text and checks it against every state it can reach. It also runs from the command line. A declaration is a short block in an ASCII notation with sets, maps, quantifiers, and a trailing apostrophe for the after-state of a variable. An excerpt from the thesis’s vending machine, including the lines that say what each phrase in a scenario asserts:
schema VendingMachine
state inventory, currentMoneyValue
const snacks = set_of("cookie", "chocolate", "chips")
const price = map_of("cookie", 5, "chocolate", 7, "chips", 3)
const maxMoney = 20
init inventory = map_of("cookie", 2, "chocolate", 1, "chips", 0)
init currentMoneyValue = 0
domain item = snacks
invariant currentMoneyValue >= 0 and currentMoneyValue <= maxMoney
operation BuyVM(item)
require item in snacks
require map_get_or(inventory, item, 0) > 0
require currentMoneyValue >= map_get(price, item)
ensure currentMoneyValue' == currentMoneyValue - map_get(price, item)
ensure inventory' == map_put(inventory, item, map_get(inventory, item) - 1)
end
fact {item} is in stock => map_get_or(inventory, item, 0) > 0
fact Bowen inserts {amount} dollars => currentMoneyValue > 0
fact Bowen selects {item} => item in snacks
goal buy a snack with money => BuyVM
Each clause is kept as a parsed tree instead of a compiled function, so a checker can read it: which clause failed, which state variables it mentions, and what the operation does to the state. Three checks follow.
Exploration. Starting from the initial state, the explorer applies every operation to every combination of values in the input domains the schema declares. It checks the invariants and postconditions in each state it reaches, and stops at a violation with the shortest sequence of operations that leads there. A LibraryLoans variant whose ReturnBook forgets to lower a member’s loan count is caught after two steps, a borrow and a return:the shortest trace is the smallest counterexample
ReturnBook reaches a state that breaks an invariant: forall m in dom(loanCount) : map_get(loanCount, m) == map_count_val(borrowedBy, m)
step 1: BorrowBook with {"Dune", "A. Diallo"}
step 2: ReturnBook with {"Dune"}
A person would have to write the scenario that reaches that state. The explorer needs none. Exhaustive here means every state reachable through the declared domains and no more, and the result says so when a search stops at its bound before finishing.
The feature check. The property in Liu’s process is that in every state where the predicates derived from a scenario’s Given and When clauses do not all hold, the goal-related operation is not enabled2. The fact and goal lines above are the analyst’s part, and the explorer’s states supply the “every state”. On listing 15 of the thesis the vending machine comes back consistent:
Goal: buy a snack with money (goal-related operation: BuyVM)
Schema states explored: 126, exhaustive: true
Scenario (positive): Buying a cookie using money when the cookie is in stock
realisable: reached by 1 step(s), then the operation runs with {"cookie"}
outcome not checked, no meaning is declared for: the vending machine dispenses a cookie
Consistent: every fact is enforced by BuyVM, and each scenario holds when the schema runs it.
Remove the requirement that enough money has been inserted, and clamp the balance at zero so the schema stays internally consistent, and the check reports the missing requirement with the state where it matters. Here that is the initial state, where a snack can be bought with no money inserted:
Requirement missing from the schema: "Bowen inserts {amount} dollars" (currentMoneyValue > 0)
BuyVM is enabled with inputs {"cookie"} in a state reached by 0 step(s), where that fact is false
Inconsistent: the feature and the schema disagree.
The Then step of the same scenario — “the vending machine dispenses a cookie” — is reported too, in both runs: outcome not checked, no meaning is declared for: the vending machine dispenses a cookie. Nothing in the schema says what dispensing means, so the claim is neither confirmed nor contradicted; it is named as unchecked rather than passed over in silence.
Realisability. The line beginning “realisable” comes from a check that is our own addition to Liu’s. Each positive scenario needs some reachable state that satisfies its facts and lets the operation run. A scenario that no reachable state satisfies, or one whose facts hold in reachable states where the operation never runs, is reported as such.
The property checked is Liu’s. Checking it by exploring the states of a schema declared in PatLang, the realisability check, and the refusal to analyse a schema that breaks its own invariants are our additions. Two limits carry over or are new. An ensure clause of the form a' == expr defines how a variable changes, and a variable with no such clause is taken as unchanged, which is a stronger reading than Z’s. A postcondition that only constrains a variable, such as n' > n, leaves the after-state open, and the schema is refused instead of explored on a guess.
Outcomes, Negative Scenarios, and Sequences
Three things the thesis leaves out2 are checked here, each by running the schema rather than trusting a scenario’s own arithmetic.
Outcomes. A then template => expr line states what a phrase claims about an operation’s result, in terms of the before-state, the after-state (primed), and the phrase’s own captures. Rather than comparing a scenario’s claimed after-state to the schema’s postcondition — which only catches a claim the postcondition itself disagrees with — the schema now computes the after-state by running the operation, and the claim is checked against what was actually computed. A claim that matches how a person hand-writing the scenario got the arithmetic wrong, but wrote a postcondition that happens to agree, is caught here where it was not before.
Negative scenarios. A Then phrase matching a negative or reject declaration marks a scenario as a rejection. It is no longer skipped: some reachable state must satisfy the scenario’s facts and disable the goal operation, and a reject phrase names the exact require clause expected to be the one that fails. A rejection blamed on the wrong clause, or one the schema never actually makes, is reported by name.
Sequences. A do template => Operation(arg, ...) line lets a When phrase run an operation instead of only asserting a fact about it. A scenario built from these is executed from every reachable state that satisfies its Given facts, one step at a time, checking each then claim against the state that step actually produced. The order events happen in is checked by happening, not modelled separately.
Refinement FoundationalKnowledge that endures for decades — core principles
Liu’s thesis raises, and leaves open, whether consistency between a behavioural specification and a Z specification survives the Z specification’s own refinement2. A concrete schema here can declare refines Abstract and one retrieve line per abstract state variable, defining it in terms of the concrete state. The check is the standard forward-simulation obligation, run over every state the concrete schema reaches: the concrete initial state retrieves to the abstract one, every retrieved state satisfies the abstract invariants, the concrete operation is enabled wherever the abstract one is (it may be enabled more often — weakening a precondition is a valid refinement), and where both are enabled, the retrieved concrete outcome equals the abstract one. A LibraryLoans variant that adds a maintained loan count refines the simple schema; the same variant with a three-loan limit added does not, and is reported with the operation, the clause that refuses it, and the trace to the state where it matters.
Once a refinement is confirmed, a feature already checked against the abstract schema can be checked against the concrete one, with each abstract fact and outcome read through the retrieve function. A concrete vending machine that adds a sales counter, correctly refining the original, keeps the feature consistent. The same concrete machine with the money-check weakened is still a valid refinement — weakening a precondition always is — but the feature is no longer consistent with it, and the missing requirement is reported by name. Confirming the refinement and confirming the feature survives it are two different checks, run separately, because the first says nothing about the second.
Inferring a Schema from a Feature Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
The boundary this section describes, stated first in the same Given/When/Then shape a feature file itself uses — a summary of the capability rather than a scenario to check, so it is shown, not run:
Feature: What BDD-to-Z inference can and cannot do
Scenario: Vocabulary alone
Given a feature that quotes or numbers the values that vary between scenarios
When it is read by zi_infer_skeleton
Then fact, do and then templates are inferred
But require and ensure stay the placeholder true
Scenario: A closed vocabulary of quantities and arithmetic
Given a feature whose Given and Then lines also say things like "contains at least N X" or "'s balance is decreased by N"
When it is read by zi_infer_skeleton_auto
Then the state variables those phrases name, and real require/ensure clauses for them, are inferred too
Scenario: Ordinary prose is never guessed at
Given a feature with no recognised phrasing, however carefully written
When it is read by either function
Then no state variable, and no require or ensure clause, is invented for it
And the phrase is left exactly as a vacuous fact or then line, unchanged
Scenario: What is never inferred, whatever the feature says
Given any feature, however precisely its scenarios are written
When it is read by zi_infer_skeleton or zi_infer_skeleton_auto
Then no init value and no operation input domain is ever produced
And the skeleton is refused when explored, not guessed into running
Scenario: Induction proposes, it never asserts
Given a schema whose facts are already filled in, and states classed accepted or rejected
When zi_suggest_requires runs
Then an already-declared fact that discriminates between them is proposed as a candidate require clause
But it is never written into the schema, or treated as checked, automatically
The direction above assumes a schema already exists. Often it doesn’t: a project accumulates Gherkin scenarios first, and the formal side never gets written. zi_infer_skeleton reads the vocabulary out of a feature and writes a schema skeleton from it, reversing the same quoting convention the fact/then/do templates above already use: a quoted span or a bare number in a Given/When/Then line generalises to a placeholder, and plain unquoted word variation does not, since deciding that “cookie is in stock” and “chips are in stock” share a template is a judgement a person has to make, not a pattern a parser can find. Given and Then lines become fact and then templates. A scenario’s When line, together with any And/But that follow it, is joined into one combined action before being generalised, so a feature that spreads one operation’s inputs across several lines — as the thesis’s own vending-machine scenario does — still infers one action of the right arity rather than several mismatched ones.
What the skeleton cannot contain is anything Gherkin text doesn’t say: no state variable, no domain, no real precondition or effect. The skeleton declares none of those, states its operation’s require and ensure clauses as the placeholder true, and leaves a TODO comment where a person’s judgement is needed, rather than guessing at any of it. It is still real, loadable PatLang the moment it is written — zs_load accepts it — because a vacuous clause and a missing declaration are both syntactically complete, only semantically empty. Run on the thesis’s own vending-machine feature, with no hand-written schema anywhere in the loop:
Output (click to expand)
schema BuyASnackWithMoneySchema
# TODO: this schema has no state yet -- Gherkin text names no
# state model, so none was guessed at. Add one, e.g.:
# state someVariable
# init someVariable = ...
operation BuyASnackWithMoney(p1)
# TODO: replace with the real precondition(s) for BuyASnackWithMoney.
require true
# TODO: replace with the real effect(s)/postcondition(s).
ensure true
end
fact cookie is in stock => true
do Bowen inserts {p1} dollars and Bowen selects cookie => BuyASnackWithMoney(p1)
then the vending machine dispenses a cookie => true
goal buy a snack with money => BuyASnackWithMoney
--- checking that the skeleton is valid, loadable PatLang ---
Loads cleanly as schema BuyASnackWithMoneySchema. Its require/ensure clauses are still the placeholder "true" above -- add real state, domains and clauses before exploring it.Comparing this against the hand-written VendingMachine schema shown earlier makes the design’s real boundary visible. Its author judged that both the dollar amount and the item name vary between scenarios, writing Bowen inserts {amount} dollars and Bowen selects {item}. The inferred skeleton catches the dollar amount, a bare number, as Bowen inserts {p1} dollars. It leaves cookie as fixed text, because the feature never quotes it and never writes a second scenario that selects a different snack — nothing in the wording signals that the word varies. A feature that does vary an unquoted value across scenarios still needs a person to notice and generalise it by hand. Exploring the skeleton as it stands reports missing_domain rather than a state count: the same honest refusal a hand-written schema with no declared domain gets, not a failure specific to inference.
What the skeleton cannot contain, above, is anything Gherkin text doesn’t say — but “doesn’t say” has its own boundary. A schema’s require/ensure clauses stay the placeholder true because free English carries no fixed formal meaning: “the cookie is nice” and “the cookie is in stock” are indistinguishable to a parser. A small, closed table of quantity and arithmetic phrasings is a different case — “contains at least N X”, “’s balance is decreased by N”, and a handful of others each have one fixed, unambiguous formal meaning, the same way a quoted span or a bare number does. zi_match_require/zi_match_ensure recognise exactly those phrasings, and only ever produce a clause when the phrase’s own subject or field literally names one of the schema’s already-declared state variables — never against arbitrary prose, and never by guessing which variable was meant.
A feature written this precisely has largely already written the formal clause in English, so this is closer to compiling a small controlled vocabulary than to reading arbitrary BDD — a real capability, conditional on that discipline, not a general upgrade to Phase 1’s own guarantee. Naming what three of the original scenario’s phrases assert, rather than leaving them as plain facts:
And vending machine inventory contains at least one cookie
And Bowen's balance is at least 10 dollars
...
And the vending machine inventory subtracts a cookie
And Bowen's balance is decreased by 10 dollars
gives real preconditions and effects, and infers the state line itself through the same table — read for the variable’s name rather than for a clause. zi_infer_state_vars proposes inventory because a require-shaped phrase names it and balance because a scalar one does, never from an unrelated noun like the item or the actor; zi_infer_skeleton_auto chains state inference and clause inference together, so nothing below was supplied by hand beyond the feature’s own text:
Output (click to expand)
schema BuyASnackWithMoneySchema
state inventory, balance
# TODO: no phrase in the feature says what these start at.
# init inventory = ...
operation BuyASnackWithMoney(p1)
require map_get_or(inventory, "cookie", 0) >= 1
require balance >= 10
ensure inventory' == map_put(inventory, "cookie", map_get(inventory, "cookie") - 1)
ensure balance' == balance - 10
end
fact cookie is in stock => true
fact vending machine inventory contains at least one cookie => true
fact Bowen's balance is at least {p1} dollars => true
do Bowen inserts {p1} dollars and Bowen selects cookie => BuyASnackWithMoney(p1)
then the vending machine dispenses a cookie => true
then the vending machine inventory subtracts a cookie => true
then Bowen's balance is decreased by {p1} dollars => true
goal buy a snack with money => BuyASnackWithMoney
--- checking that the skeleton is valid, loadable PatLang ---
Loads cleanly as schema BuyASnackWithMoneySchema. State (inventory, balance) and the require/ensure clauses above came from the feature's own wording -- add init values and an input domain before exploring it.A feature with no recognised phrase at all — the plain version run earlier — gets Phase 1’s output back completely unchanged, confirmed by a regression test rather than left to assumption. Hand-completing this skeleton with init values and a domain for p1 (still nobody’s job but a person’s: no phrase says what a snack machine starts with, or what amounts a buyer might insert) makes it explorable: two reachable states, one transition, a cookie bought.
A second, bounded step goes further, but only once a person has filled a schema’s facts in with real meaning. For an operation with some scenarios classed as accepted and others rejected, zi_suggest_requires proposes an already-declared fact as a candidate require clause whenever it holds in every accepted state and fails in at least one rejected one. It searches nothing beyond the vocabulary already written down — this is not expression synthesis — and every result carries its own support and counter-example counts rather than being asserted as checked:
require amount is small (holds in 2/2 accepted states; excludes 1/1 rejected states)
A fact merely correlating with acceptance is not the same claim as a fact a schema’s author intended as a precondition, so a suggestion here is exactly that: a hypothesis for a person to confirm, or reject, before it becomes a require clause the schema is actually checked against. Deriving accepted and rejected states automatically from a checked feature, rather than supplying them by hand as above, is the natural next integration and isn’t built yet. Full source: self_hosting/lib/zs_infer.patlang. The generated examples above are checked into self_hosting/examples/zs/vending_machine_inferred.zschema and self_hosting/examples/zs/priced_vending_machine_inferred.zschema, each alongside the feature it came from.
Program Flow Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
A declaration is loaded once and kept as parsed clauses. The explorer, the feature analysis, the generator, and the refinement check all read from that same entry.
state, invariants, operations"] --> L["zs_load
parse each clause into a tree"] L --> R[("Registry entry")] R --> E["zs_explore
breadth-first search of reachable states"] E -->|"invariant or postcondition broken"| V["Violation:
shortest sequence of operations"] E -->|"every state checked"| OK["ok, or ok_bounded
if the state limit was hit"] F["Feature text
Given / When / Then"] --> FP["zf_parse_feature
goal, scenarios, steps"] FP --> MF["match each step to a fact
declared in the schema"] R --> MF E -->|"reachable states"| AN MF --> AN["zs_analyse_feature
outcomes, negative scenarios,
sequences, realisability, necessity"] AN --> RES["consistent or inconsistent,
with findings"] R --> GEN["zs_generate_feature
one scenario per witnessed clause"] GEN -.->|"fed back in"| F R --> REF["zs_refines
forward-simulation obligations"] REF --> RRES["a concrete schema either
refines the abstract one, or doesn't"] F --> INF["zi_infer_skeleton_auto
vocabulary, plus a closed table of
quantity/arithmetic phrasings"] INF --> SK[("Schema skeleton
state and require/ensure are real where a
phrase names them, else placeholder true")] SK -.->|"a person adds init values,
an input domain, and any
clause still a placeholder"| L style V fill:#F5ECDB style RES fill:#E8F5E9 style RRES fill:#E8F5E9 style SK fill:#F5ECDB
Inside the explorer, each state is expanded by trying every operation with every combination of the declared input values. A clock, not a state count, decides when the search reports progress, so a state with thousands of transitions cannot delay a report.
next combination of inputs"] OP --> PRE{"all require
clauses hold?"} PRE -->|no| OP PRE -->|yes| EFF["compute the after-state
from the ensure equations"] EFF --> CHK{"ensure clauses and
invariants hold?"} CHK -->|no| BAD["stop and report
the trace to this step"] CHK -->|yes| SEEN{"state already seen?
canonical key"} SEEN -->|yes| OP SEEN -->|no| ADD["queue it and record
the step that reached it"] ADD --> OP OP -->|"all combinations tried"| Q Q -->|"queue empty"| DONE["every reachable state checked"] T["clock tick, every 120 ms
progress line, status, quit"] -.-> OP style BAD fill:#F5ECDB style DONE fill:#E8F5E9
Try It
The panel below runs the same libraries in this page, in a background worker so the page stays responsive. Beyond exploring a schema and checking a feature against it, the dropdown includes an outcome-checking example, a negative scenario, a multi-step sequence, the schema-to-Gherkin generator, a data refinement, and inferring a schema skeleton (including real preconditions, effects and state, where the feature names them precisely enough) from a feature with no schema at all — each run with the same code as the command-line tools. Edit the schema or the feature and run it again; Cancel stops the worker.
(not run yet)
make a function called object_delete takes name returns done
return true
end
make a function called list_copy takes l returns out
let out = []
let i = 0
let n = to_num(list_len(l))
while i < n do
let out = list_push(out, l[i])
let i = i + 1
end
return out
end
make a function called list_appended takes l, v returns out
let out = list_copy(l)
return list_push(out, v)
end
make a function called pset_new returns s
return []
end
make a function called pset_contains takes s, item returns found
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] == item then
return true
end
let i = i + 1
end
return false
end
make a function called pset_add takes s, item returns s2
if pset_contains(s, item) then
return s
end
return list_appended(s, item)
end
make a function called pset_remove takes s, item returns s2
let out = []
let i = 0
let n = to_num(list_len(s))
while i < n do
if s[i] != item then
let out = list_push(out, s[i])
end
let i = i + 1
end
return out
end
make a function called pset_size takes s returns n
return to_num(list_len(s))
end
make a function called pset_to_list takes s returns items
return s
end
make a function called pset_equal takes a, b returns eq
if pset_size(a) != pset_size(b) then
return false
end
let i = 0
let n = to_num(list_len(a))
while i < n do
if pset_contains(b, a[i]) == false then
return false
end
let i = i + 1
end
return true
end
make a function called pmap_new returns m
return []
end
make a function called pmap_has takes m, key returns found
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return true
end
let i = i + 1
end
return false
end
make a function called pmap_get takes m, key returns value
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
return m[i][1]
end
let i = i + 1
end
return []
end
make a function called pmap_put takes m, key, value returns m2
let out = []
let found = false
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] == key then
let out = list_push(out, [key, value])
let found = true
else
let out = list_push(out, m[i])
end
let i = i + 1
end
if found == false then
let out = list_push(out, [key, value])
end
return out
end
make a function called pmap_remove takes m, key returns m2
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
if m[i][0] != key then
let out = list_push(out, m[i])
end
let i = i + 1
end
return out
end
make a function called pmap_keys takes m returns keys
let out = []
let i = 0
let n = to_num(list_len(m))
while i < n do
let out = list_push(out, m[i][0])
let i = i + 1
end
return out
end
make a function called pmap_size takes m returns n
return to_num(list_len(m))
end
make a function called sm_parse_template takes template returns parsed
let h = str_intern(template)
let n = sc_len(h)
let literals = []
let names = []
let cur = sb_new()
let i = 0
while i < n do
let c = sc_code(h, i)
if c == 123 then
let j = i + 1
let nm = sb_new()
while (j < n) and (sc_code(h, j) != 125) do
sb_push(nm, sc_char(h, j))
let j = j + 1
end
let literals = list_push(literals, sb_str(cur))
let cur = sb_new()
let names = list_push(names, sb_str(nm))
let i = j + 1
else
sb_push(cur, sc_char(h, i))
let i = i + 1
end
end
let literals = list_push(literals, sb_str(cur))
return [literals, names]
end
make a function called sm_at takes h, hn, lit, at returns found
let hl = str_intern(lit)
let ln = sc_len(hl)
if at < 0 then
return false
end
if at + ln > hn then
return false
end
let k = 0
while k < ln do
if sc_code(h, at + k) != sc_code(hl, k) then
return false
end
let k = k + 1
end
return true
end
make a function called sm_find takes h, hn, lit, from returns idx
let ln = sc_len(str_intern(lit))
let i = from
while i + ln <= hn do
if sm_at(h, hn, lit, i) then
return i
end
let i = i + 1
end
return -1
end
make a function called sm_match takes template, text returns result
let parsed = sm_parse_template(template)
let lits = parsed[0]
let names = parsed[1]
let nn = to_num(list_len(names))
let ht = str_intern(text)
let tn = sc_len(ht)
let lead = lits[0]
let lead_len = sc_len(str_intern(lead))
if sm_at(ht, tn, lead, 0) == false then
return ["no", []]
end
if nn == 0 then
if lead_len == tn then
return ["ok", []]
end
return ["no", []]
end
let pos = lead_len
let caps = []
let k = 0
while k < nn do
let lit = lits[k + 1]
let lit_len = sc_len(str_intern(lit))
let cap = ""
if k == nn - 1 then
let stop = tn - lit_len
if stop < pos then
return ["no", []]
end
if sm_at(ht, tn, lit, stop) == false then
return ["no", []]
end
let cap = sc_substr(ht, pos, stop - pos)
let pos = tn
else
if lit_len == 0 then
return ["no", []]
end
let idx = sm_find(ht, tn, lit, pos)
if idx < 0 then
return ["no", []]
end
let cap = sc_substr(ht, pos, idx - pos)
let pos = idx + lit_len
end
if cap == "" then
return ["no", []]
end
let caps = list_push(caps, [names[k], cap])
let k = k + 1
end
return ["ok", caps]
end
make a function called sm_fill takes template, bindings returns text
let parsed = sm_parse_template(template)
let lits = parsed[0]
let names = parsed[1]
let out = sb_new()
sb_push(out, lits[0])
let k = 0
let kn = to_num(list_len(names))
while k < kn do
let found = false
let b = 0
let bn = to_num(list_len(bindings))
while b < bn do
if bindings[b][0] == names[k] then
sb_push(out, "" + bindings[b][1])
let found = true
end
let b = b + 1
end
if found == false then
sb_push(out, "{" + names[k] + "}")
end
sb_push(out, lits[k + 1])
let k = k + 1
end
return sb_str(out)
end
make a function called schema_registry returns registry
let existing = get("__vars", "schema_bdd_registry_obj")
if existing then
return existing
end
let registry = new("Dict", "schema_bdd_registry")
set_var("schema_bdd_registry_obj", registry)
return registry
end
make a function called schema_define takes name, state_var_names, invariant_fn returns done
send(schema_registry(), "set", name + "__schema", [state_var_names, invariant_fn])
return true
end
make a function called schema_lookup takes name returns entry
return get(schema_registry(), name + "__schema")
end
make a function called schema_operation takes schema_name, op_name, input_names, require_fn, ensure_fn returns done
send(schema_registry(), "set", schema_name + "::" + op_name + "__op", [schema_name, input_names, require_fn, ensure_fn])
return true
end
make a function called schema_lookup_operation takes schema_name, op_name returns entry
return get(schema_registry(), schema_name + "::" + op_name + "__op")
end
make a function called schema_bind_state takes schema, var_name, phase, value returns done
set_var(schema + "__" + var_name + "__" + phase, value)
return true
end
make a function called schema_bind_input takes schema, op, param_name, value returns done
set_var(schema + "__" + op + "__in__" + param_name, value)
return true
end
make a function called schema_state_values takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__before"))
let i = i + 1
end
return out
end
make a function called schema_state_values_after takes schema, state_var_names returns values
let out = []
let i = 0
let n = to_num(list_len(state_var_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + state_var_names[i] + "__after"))
let i = i + 1
end
return out
end
make a function called schema_input_values takes schema, op, input_names returns values
let out = []
let i = 0
let n = to_num(list_len(input_names))
while i < n do
let out = list_push(out, get("__vars", schema + "__" + op + "__in__" + input_names[i]))
let i = i + 1
end
return out
end
make a function called schema_check_values takes schema_name, op_name, state_before, inputs, state_after returns diagnosis
let schema_entry = schema_lookup(schema_name)
let invariant_fn = schema_entry[1]
let op_entry = schema_lookup_operation(schema_name, op_name)
let require_fn = op_entry[2]
let ensure_fn = op_entry[3]
if apply(invariant_fn, state_before) == false then
return ["invariant_violated_before", [schema_name, state_before]]
end
if apply(require_fn, state_before, inputs) == false then
return ["precondition_violated", [op_name, state_before, inputs]]
end
if apply(invariant_fn, state_after) == false then
return ["invariant_violated_after", [schema_name, state_after]]
end
if apply(ensure_fn, state_before, state_after, inputs) == false then
return ["postcondition_violated", [op_name, state_before, state_after, inputs]]
end
return ["ok", [op_name, state_before, state_after, inputs]]
end
make a function called schema_check takes schema_name, op_name returns diagnosis
let schema_entry = schema_lookup(schema_name)
let state_var_names = schema_entry[0]
let op_entry = schema_lookup_operation(schema_name, op_name)
let input_names = op_entry[1]
let state_before = schema_state_values(schema_name, state_var_names)
let inputs = schema_input_values(schema_name, op_name, input_names)
let state_after = schema_state_values_after(schema_name, state_var_names)
return schema_check_values(schema_name, op_name, state_before, inputs, state_after)
end
make a function called schema_format_diagnosis takes schema_name, op_name, diagnosis returns questions
let tag = diagnosis[0]
let payload = diagnosis[1]
if tag == "invariant_violated_before" then
return ["This scenario's Given already leaves " + schema_name + " in a state that violates its own invariant, before " + op_name + " is even checked. Is the Given wrong, or is the invariant too strict?"]
end
if tag == "precondition_violated" then
return ["Scenario claims " + op_name + " can run from this Given, but " + op_name + "'s own precondition returned false for these inputs. Is the Given wrong, or is " + op_name + "'s precondition too strict -- or is this scenario meant to test a rejection path, which needs its own operation schema rather than " + op_name + "'s?"]
end
if tag == "invariant_violated_after" then
return ["Applying " + op_name + " produces a state that violates " + schema_name + "'s invariant. Is the Then clause's claimed resulting state wrong, or does " + op_name + " need a stronger precondition to rule this case out?"]
end
if tag == "postcondition_violated" then
return ["The before/after state this scenario claims for " + op_name + " does not satisfy its own postcondition. Is the Then clause's claimed resulting state wrong, or is " + op_name + "'s postcondition wrong?"]
end
return []
end
make a function called schema_register_synthesis_check takes func_name, schema_name, op_name, harness_fn returns done
send(schema_registry(), "set", func_name + "__synthesis_check", [schema_name, op_name, harness_fn])
return true
end
make a function called schema_lookup_synthesis_check takes func_name returns entry
return get(schema_registry(), func_name + "__synthesis_check")
end
make a function called schema_run_synthesis_check takes func_name returns diagnosis
let hook = schema_lookup_synthesis_check(func_name)
let schema_name = hook[0]
let op_name = hook[1]
let harness_fn = hook[2]
let triple = apply(harness_fn, func_name)
return schema_check_values(schema_name, op_name, triple[0], triple[1], triple[2])
end
make a function called ze_rt_error returns msg
let v = get("__vars", "ze_rt_error")
if type_of(v) != "string" then
return ""
end
return v
end
make a function called ze_rt_clear returns done
set_var("ze_rt_error", "")
return true
end
make a function called ze_rt_fail takes msg returns done
if ze_rt_error() == "" then
set_var("ze_rt_error", msg)
end
return true
end
make a function called ze_parse_error returns msg
let v = get("__vars", "ze_error")
if type_of(v) != "string" then
return ""
end
return v
end
make a function called ze_parse_fail takes msg returns done
if ze_parse_error() == "" then
set_var("ze_error", msg)
end
return true
end
make a function called ze_is_alpha takes c returns r
return ((c >= 65) and (c <= 90)) or ((c >= 97) and (c <= 122)) or (c == 95)
end
make a function called ze_is_digit takes c returns r
return (c >= 48) and (c <= 57)
end
make a function called ze_is_keyword takes w returns r
if w == "and" then
return true
end
if w == "or" then
return true
end
if w == "not" then
return true
end
if w == "implies" then
return true
end
if w == "in" then
return true
end
if w == "forall" then
return true
end
if w == "exists" then
return true
end
if w == "true" then
return true
end
if w == "false" then
return true
end
return false
end
make a function called ze_tokenize takes text returns toks
let h = str_intern(text)
let n = sc_len(h)
let toks = []
let i = 0
while i < n do
let c = sc_code(h, i)
if (c == 32) or (c == 9) or (c == 13) or (c == 10) then
let i = i + 1
elif ze_is_digit(c) then
let b = sb_new()
let j = i
while (j < n) and ze_is_digit(sc_code(h, j)) do
sb_push(b, sc_char(h, j))
let j = j + 1
end
let toks = list_push(toks, ["num", sb_str(b)])
let i = j
elif ze_is_alpha(c) then
let b = sb_new()
let j = i
while (j < n) and (ze_is_alpha(sc_code(h, j)) or ze_is_digit(sc_code(h, j))) do
sb_push(b, sc_char(h, j))
let j = j + 1
end
if (j < n) and (sc_code(h, j) == 39) then
sb_push(b, "'")
let j = j + 1
end
let word = sb_str(b)
if ze_is_keyword(word) then
let toks = list_push(toks, ["kw", word])
else
let toks = list_push(toks, ["id", word])
end
let i = j
elif c == 34 then
let b = sb_new()
let j = i + 1
while (j < n) and (sc_code(h, j) != 34) do
sb_push(b, sc_char(h, j))
let j = j + 1
end
if j >= n then
ze_parse_fail("unterminated string")
end
let toks = list_push(toks, ["str", sb_str(b)])
let i = j + 1
else
let nx = -1
if i + 1 < n then
let nx = sc_code(h, i + 1)
end
if (c == 61) and (nx == 61) then
let toks = list_push(toks, ["op", "=="])
let i = i + 2
elif (c == 33) and (nx == 61) then
let toks = list_push(toks, ["op", "!="])
let i = i + 2
elif (c == 60) and (nx == 61) then
let toks = list_push(toks, ["op", "<="])
let i = i + 2
elif (c == 62) and (nx == 61) then
let toks = list_push(toks, ["op", ">="])
let i = i + 2
elif (c == 60) or (c == 62) or (c == 40) or (c == 41) or (c == 44) or (c == 58) or (c == 43) or (c == 45) or (c == 42) or (c == 37) then
let toks = list_push(toks, ["op", sc_char(h, i)])
let i = i + 1
else
ze_parse_fail("unexpected character " + sc_char(h, i))
let i = i + 1
end
end
end
return toks
end
make a function called ze_tok_is takes toks, pos, kind, text returns r
if pos >= to_num(list_len(toks)) then
return false
end
let t = toks[pos]
if t[0] != kind then
return false
end
return t[1] == text
end
make a function called ze_parse_implies takes toks, pos returns r
let a = ze_parse_or(toks, pos)
let p = a[1]
if ze_tok_is(toks, p, "kw", "implies") then
let b = ze_parse_implies(toks, p + 1)
return [["bin", "implies", a[0], b[0]], b[1]]
end
return a
end
make a function called ze_parse_or takes toks, pos returns r
let a = ze_parse_and(toks, pos)
let node = a[0]
let p = a[1]
while ze_tok_is(toks, p, "kw", "or") do
let b = ze_parse_and(toks, p + 1)
let node = ["bin", "or", node, b[0]]
let p = b[1]
end
return [node, p]
end
make a function called ze_parse_and takes toks, pos returns r
let a = ze_parse_not(toks, pos)
let node = a[0]
let p = a[1]
while ze_tok_is(toks, p, "kw", "and") do
let b = ze_parse_not(toks, p + 1)
let node = ["bin", "and", node, b[0]]
let p = b[1]
end
return [node, p]
end
make a function called ze_parse_not takes toks, pos returns r
if ze_tok_is(toks, pos, "kw", "not") then
let a = ze_parse_not(toks, pos + 1)
return [["not", a[0]], a[1]]
end
return ze_parse_cmp(toks, pos)
end
make a function called ze_cmp_op takes toks, pos returns op
if pos >= to_num(list_len(toks)) then
return ""
end
let t = toks[pos]
if t[0] == "op" then
if (t[1] == "==") or (t[1] == "!=") or (t[1] == "<") or (t[1] == "<=") or (t[1] == ">") or (t[1] == ">=") then
return t[1]
end
end
if (t[0] == "kw") and (t[1] == "in") then
return "in"
end
return ""
end
make a function called ze_parse_cmp takes toks, pos returns r
let a = ze_parse_add(toks, pos)
let op = ze_cmp_op(toks, a[1])
if op != "" then
let b = ze_parse_add(toks, a[1] + 1)
return [["bin", op, a[0], b[0]], b[1]]
end
return a
end
make a function called ze_parse_add takes toks, pos returns r
let a = ze_parse_mul(toks, pos)
let node = a[0]
let p = a[1]
let going = true
while going do
if ze_tok_is(toks, p, "op", "+") then
let b = ze_parse_mul(toks, p + 1)
let node = ["bin", "+", node, b[0]]
let p = b[1]
elif ze_tok_is(toks, p, "op", "-") then
let b = ze_parse_mul(toks, p + 1)
let node = ["bin", "-", node, b[0]]
let p = b[1]
else
let going = false
end
end
return [node, p]
end
make a function called ze_parse_mul takes toks, pos returns r
let a = ze_parse_unary(toks, pos)
let node = a[0]
let p = a[1]
let going = true
while going do
if ze_tok_is(toks, p, "op", "*") then
let b = ze_parse_unary(toks, p + 1)
let node = ["bin", "*", node, b[0]]
let p = b[1]
elif ze_tok_is(toks, p, "op", "%") then
let b = ze_parse_unary(toks, p + 1)
let node = ["bin", "%", node, b[0]]
let p = b[1]
else
let going = false
end
end
return [node, p]
end
make a function called ze_parse_unary takes toks, pos returns r
if ze_tok_is(toks, pos, "op", "-") then
let a = ze_parse_unary(toks, pos + 1)
return [["neg", a[0]], a[1]]
end
return ze_parse_atom(toks, pos)
end
make a function called ze_parse_atom takes toks, pos returns r
let n = to_num(list_len(toks))
if pos >= n then
ze_parse_fail("unexpected end of clause")
return [["num", 0], pos]
end
let t = toks[pos]
let kind = t[0]
let text = t[1]
if kind == "num" then
return [["num", to_num(text)], pos + 1]
end
if kind == "str" then
return [["str", text], pos + 1]
end
if kind == "kw" then
if text == "true" then
return [["bool", true], pos + 1]
end
if text == "false" then
return [["bool", false], pos + 1]
end
if (text == "forall") or (text == "exists") then
if (pos + 1 >= n) or (toks[pos + 1][0] != "id") then
ze_parse_fail(text + " needs a variable name")
return [["num", 0], pos + 1]
end
let v = toks[pos + 1][1]
if ze_tok_is(toks, pos + 2, "kw", "in") == false then
ze_parse_fail(text + " " + v + " needs 'in'")
return [["num", 0], pos + 2]
end
let dom = ze_parse_add(toks, pos + 3)
if ze_tok_is(toks, dom[1], "op", ":") == false then
ze_parse_fail(text + " " + v + " needs ':' before its body")
return [["num", 0], dom[1]]
end
let body = ze_parse_implies(toks, dom[1] + 1)
return [[text, v, dom[0], body[0]], body[1]]
end
ze_parse_fail("unexpected keyword " + text)
return [["num", 0], pos + 1]
end
if kind == "id" then
if ze_tok_is(toks, pos + 1, "op", "(") then
let args = []
let p = pos + 2
if ze_tok_is(toks, p, "op", ")") then
return [["call", text, args], p + 1]
end
let going = true
while going do
let a = ze_parse_implies(toks, p)
let args = list_push(args, a[0])
let p = a[1]
if ze_tok_is(toks, p, "op", ",") then
let p = p + 1
else
let going = false
end
end
if ze_tok_is(toks, p, "op", ")") == false then
ze_parse_fail("unclosed call to " + text)
return [["call", text, args], p]
end
return [["call", text, args], p + 1]
end
return [["var", text], pos + 1]
end
if (kind == "op") and (text == "(") then
let a = ze_parse_implies(toks, pos + 1)
if ze_tok_is(toks, a[1], "op", ")") == false then
ze_parse_fail("missing )")
return a
end
return [a[0], a[1] + 1]
end
ze_parse_fail("unexpected " + text)
return [["num", 0], pos + 1]
end
make a function called ze_parse takes text returns result
set_var("ze_error", "")
let toks = ze_tokenize(text)
if ze_parse_error() != "" then
return ["error", ze_parse_error()]
end
if to_num(list_len(toks)) == 0 then
return ["error", "empty clause"]
end
let r = ze_parse_implies(toks, 0)
if ze_parse_error() != "" then
return ["error", ze_parse_error()]
end
if r[1] < to_num(list_len(toks)) then
return ["error", "unexpected " + toks[r[1]][1] + " after the end of the clause"]
end
return ["ok", r[0]]
end
make a function called ze_names_has takes names, name returns found
let i = 0
let n = to_num(list_len(names))
while i < n do
if names[i] == name then
return true
end
let i = i + 1
end
return false
end
make a function called ze_fv takes ast, bound, acc returns out
let k = ast[0]
if k == "var" then
if ze_names_has(bound, ast[1]) == false then
if ze_names_has(acc, ast[1]) == false then
return list_push(acc, ast[1])
end
end
return acc
end
if (k == "not") or (k == "neg") then
return ze_fv(ast[1], bound, acc)
end
if k == "bin" then
let acc = ze_fv(ast[2], bound, acc)
return ze_fv(ast[3], bound, acc)
end
if k == "call" then
let args = ast[2]
let i = 0
let n = to_num(list_len(args))
while i < n do
let acc = ze_fv(args[i], bound, acc)
let i = i + 1
end
return acc
end
if (k == "forall") or (k == "exists") then
let acc = ze_fv(ast[2], bound, acc)
return ze_fv(ast[3], list_appended(bound, ast[1]), acc)
end
return acc
end
make a function called ze_free_vars takes ast returns names
return ze_fv(ast, [], [])
end
make a function called ze_subst takes ast, pairs, bound returns out
let k = ast[0]
if k == "var" then
if ze_names_has(bound, ast[1]) then
return ast
end
let i = 0
let n = to_num(list_len(pairs))
while i < n do
if pairs[i][0] == ast[1] then
return pairs[i][1]
end
let i = i + 1
end
return ast
end
if (k == "not") or (k == "neg") then
return [k, ze_subst(ast[1], pairs, bound)]
end
if k == "bin" then
return ["bin", ast[1], ze_subst(ast[2], pairs, bound), ze_subst(ast[3], pairs, bound)]
end
if k == "call" then
let args = []
let j = 0
let jn = to_num(list_len(ast[2]))
while j < jn do
let args = list_push(args, ze_subst(ast[2][j], pairs, bound))
let j = j + 1
end
return ["call", ast[1], args]
end
if (k == "forall") or (k == "exists") then
return [k, ast[1], ze_subst(ast[2], pairs, bound), ze_subst(ast[3], pairs, list_appended(bound, ast[1]))]
end
return ast
end
make a function called ze_is_atom takes v returns r
return type_of(v) != "list"
end
make a function called ze_is_pair takes e returns r
if type_of(e) != "list" then
return false
end
if to_num(list_len(e)) != 2 then
return false
end
return type_of(e[0]) != "list"
end
make a function called ze_is_numeric_type takes t returns r
return (t == "int") or (t == "float") or (t == "bigint") or (t == "rational") or (t == "complex")
end
make a function called ze_val_eq takes a, b returns r
let ta = type_of(a)
let tb = type_of(b)
if ta != tb then
if ze_is_numeric_type(ta) and ze_is_numeric_type(tb) then
return a == b
end
return false
end
if ta != "list" then
return a == b
end
let n = to_num(list_len(a))
if n != to_num(list_len(b)) then
return false
end
let i = 0
while i < n do
if ze_elem_in(a[i], b) == false then
return false
end
let i = i + 1
end
return true
end
make a function called ze_elem_in takes e, c returns found
let n = to_num(list_len(c))
let j = 0
while j < n do
if ze_elem_eq(e, c[j]) then
return true
end
let j = j + 1
end
return false
end
make a function called ze_elem_eq takes x, y returns r
if ze_is_pair(x) and ze_is_pair(y) then
return ze_val_eq(x[0], y[0]) and ze_val_eq(x[1], y[1])
end
return ze_val_eq(x, y)
end
make a function called ze_key_of takes item, ki returns k
if ki < 0 then
return item
end
return item[ki]
end
make a function called ze_msort takes items, ki returns out
let n = to_num(list_len(items))
if n <= 1 then
return items
end
let mid = (n - (n % 2)) / 2
let left = []
let right = []
let i = 0
while i < n do
if i < mid then
let left = list_push(left, items[i])
else
let right = list_push(right, items[i])
end
let i = i + 1
end
let ls = ze_msort(left, ki)
let rs = ze_msort(right, ki)
let ln = to_num(list_len(ls))
let rn = to_num(list_len(rs))
let out = []
let a = 0
let b = 0
while (a < ln) and (b < rn) do
if ze_key_of(rs[b], ki) < ze_key_of(ls[a], ki) then
let out = list_push(out, rs[b])
let b = b + 1
else
let out = list_push(out, ls[a])
let a = a + 1
end
end
while a < ln do
let out = list_push(out, ls[a])
let a = a + 1
end
while b < rn do
let out = list_push(out, rs[b])
let b = b + 1
end
return out
end
make a function called ze_atom_canon takes v returns s
let t = type_of(v)
if t == "int" then
return "i" + v
end
if t == "string" then
return "s" + v.length + ":" + v
end
if t == "bool" then
if v then
return "b1"
end
return "b0"
end
return "?" + t + ":" + v
end
make a function called ze_canon takes v returns s
if type_of(v) != "list" then
return ze_atom_canon(v)
end
let n = to_num(list_len(v))
let parts = []
let i = 0
while i < n do
let e = v[i]
if ze_is_pair(e) then
let parts = list_push(parts, "(" + ze_canon(e[0]) + "," + ze_canon(e[1]) + ")")
else
let parts = list_push(parts, ze_canon(e))
end
let i = i + 1
end
let sorted = ze_msort(parts, -1)
let b = sb_new()
sb_push(b, "{")
let k = 0
while k < n do
if k > 0 then
sb_push(b, ",")
end
sb_push(b, sorted[k])
let k = k + 1
end
sb_push(b, "}")
return sb_str(b)
end
make a function called ze_show takes v returns s
let t = type_of(v)
if t == "string" then
return "\"" + v + "\""
end
if t == "bool" then
if v then
return "true"
end
return "false"
end
if t != "list" then
return "" + v
end
let b = sb_new()
sb_push(b, "{")
let n = to_num(list_len(v))
let i = 0
while i < n do
if i > 0 then
sb_push(b, ", ")
end
if ze_is_pair(v[i]) then
sb_push(b, ze_show(v[i][0]))
sb_push(b, " -> ")
sb_push(b, ze_show(v[i][1]))
else
sb_push(b, ze_show(v[i]))
end
let i = i + 1
end
sb_push(b, "}")
return sb_str(b)
end
make a function called ze_set_of takes items returns s
let out = pset_new()
let i = 0
let n = to_num(list_len(items))
while i < n do
let out = pset_add(out, items[i])
let i = i + 1
end
return out
end
make a function called ze_set_union takes a, b returns s
let out = a
let i = 0
let n = to_num(list_len(b))
while i < n do
let out = pset_add(out, b[i])
let i = i + 1
end
return out
end
make a function called ze_set_diff takes a, b returns s
let out = pset_new()
let i = 0
let n = to_num(list_len(a))
while i < n do
if pset_contains(b, a[i]) == false then
let out = pset_add(out, a[i])
end
let i = i + 1
end
return out
end
make a function called ze_set_inter takes a, b returns s
let out = pset_new()
let i = 0
let n = to_num(list_len(a))
while i < n do
if pset_contains(b, a[i]) then
let out = pset_add(out, a[i])
end
let i = i + 1
end
return out
end
make a function called ze_subset takes a, b returns r
let i = 0
let n = to_num(list_len(a))
while i < n do
if pset_contains(b, a[i]) == false then
return false
end
let i = i + 1
end
return true
end
make a function called ze_ran takes m returns s
let out = pset_new()
let i = 0
let n = to_num(list_len(m))
while i < n do
let out = pset_add(out, m[i][1])
let i = i + 1
end
return out
end
make a function called ze_image takes m, keys returns s
let out = pset_new()
let i = 0
let n = to_num(list_len(keys))
while i < n do
if pmap_has(m, keys[i]) then
let out = pset_add(out, pmap_get(m, keys[i]))
end
let i = i + 1
end
return out
end
make a function called ze_map_count_val takes m, v returns n
let count = 0
let i = 0
let len = to_num(list_len(m))
while i < len do
if ze_val_eq(m[i][1], v) then
let count = count + 1
end
let i = i + 1
end
return count
end
make a function called ze_map_of takes items returns m
let out = pmap_new()
let i = 0
let n = to_num(list_len(items))
while i + 1 < n do
let out = pmap_put(out, items[i], items[i + 1])
let i = i + 2
end
return out
end
make a function called ze_register_fn takes name, fname returns done
let fns = get("__vars", "ze_user_fns")
if type_of(fns) != "list" then
let fns = []
end
set_var("ze_user_fns", list_appended(fns, [name, fname]))
return true
end
make a function called ze_user_fn takes name returns fname
let fns = get("__vars", "ze_user_fns")
if type_of(fns) != "list" then
return ""
end
let i = to_num(list_len(fns)) - 1
while i >= 0 do
if fns[i][0] == name then
return fns[i][1]
end
let i = i - 1
end
return ""
end
make a function called ze_call_user takes fname, args returns v
let n = to_num(list_len(args))
if n == 0 then
return apply(fname)
end
if n == 1 then
return apply(fname, args[0])
end
if n == 2 then
return apply(fname, args[0], args[1])
end
if n == 3 then
return apply(fname, args[0], args[1], args[2])
end
if n == 4 then
return apply(fname, args[0], args[1], args[2], args[3])
end
ze_rt_fail("a registered function takes at most 4 arguments: " + fname)
return false
end
make a function called ze_arity takes name, args, want returns ok
if to_num(list_len(args)) != want then
ze_rt_fail(name + " takes " + want + " argument(s), got " + to_num(list_len(args)))
return false
end
return true
end
make a function called ze_lead_lists takes name returns n
if (name == "set_union") or (name == "set_diff") or (name == "set_inter") or (name == "subset") or (name == "image") then
return 2
end
if (name == "size") or (name == "map_size") or (name == "set_add") or (name == "set_remove") or (name == "map_has") or (name == "map_get") or (name == "map_get_or") or (name == "map_put") or (name == "map_remove") or (name == "dom") or (name == "map_count_val") or (name == "ran") then
return 1
end
return 0
end
make a function called ze_lists_ok takes name, args returns ok
let lead = ze_lead_lists(name)
if lead == 0 then
return true
end
if to_num(list_len(args)) < lead then
ze_rt_fail(name + " is missing arguments")
return false
end
let i = 0
while i < lead do
if type_of(args[i]) != "list" then
ze_rt_fail(name + " needs a set or map as argument " + (i + 1) + ", got a " + type_of(args[i]))
return false
end
let i = i + 1
end
return true
end
make a function called ze_call takes name, args returns v
if ze_lists_ok(name, args) == false then
return false
end
if name == "set_of" then
return ze_set_of(args)
end
if name == "map_of" then
return ze_map_of(args)
end
if name == "empty_set" then
return pset_new()
end
if name == "empty_map" then
return pmap_new()
end
if name == "size" then
if ze_arity(name, args, 1) == false then
return false
end
return to_num(list_len(args[0]))
end
if name == "map_size" then
if ze_arity(name, args, 1) == false then
return false
end
return to_num(list_len(args[0]))
end
if name == "set_add" then
if ze_arity(name, args, 2) == false then
return false
end
return pset_add(args[0], args[1])
end
if name == "set_remove" then
if ze_arity(name, args, 2) == false then
return false
end
return pset_remove(args[0], args[1])
end
if name == "set_union" then
if ze_arity(name, args, 2) == false then
return false
end
return ze_set_union(args[0], args[1])
end
if name == "set_diff" then
if ze_arity(name, args, 2) == false then
return false
end
return ze_set_diff(args[0], args[1])
end
if name == "set_inter" then
if ze_arity(name, args, 2) == false then
return false
end
return ze_set_inter(args[0], args[1])
end
if name == "subset" then
if ze_arity(name, args, 2) == false then
return false
end
return ze_subset(args[0], args[1])
end
if name == "map_has" then
if ze_arity(name, args, 2) == false then
return false
end
return pmap_has(args[0], args[1])
end
if name == "map_get" then
if ze_arity(name, args, 2) == false then
return false
end
if pmap_has(args[0], args[1]) == false then
ze_rt_fail("map_get: key " + args[1] + " is not in the map")
return false
end
return pmap_get(args[0], args[1])
end
if name == "map_get_or" then
if ze_arity(name, args, 3) == false then
return false
end
if pmap_has(args[0], args[1]) then
return pmap_get(args[0], args[1])
end
return args[2]
end
if name == "map_put" then
if ze_arity(name, args, 3) == false then
return false
end
return pmap_put(args[0], args[1], args[2])
end
if name == "map_remove" then
if ze_arity(name, args, 2) == false then
return false
end
return pmap_remove(args[0], args[1])
end
if name == "map_count_val" then
if ze_arity(name, args, 2) == false then
return false
end
return ze_map_count_val(args[0], args[1])
end
if name == "dom" then
if ze_arity(name, args, 1) == false then
return false
end
return pmap_keys(args[0])
end
if name == "ran" then
if ze_arity(name, args, 1) == false then
return false
end
return ze_ran(args[0])
end
if name == "image" then
if ze_arity(name, args, 2) == false then
return false
end
return ze_image(args[0], args[1])
end
if name == "min" then
if ze_arity(name, args, 2) == false then
return false
end
if args[1] < args[0] then
return args[1]
end
return args[0]
end
if name == "max" then
if ze_arity(name, args, 2) == false then
return false
end
if args[1] > args[0] then
return args[1]
end
return args[0]
end
if name == "abs" then
if ze_arity(name, args, 1) == false then
return false
end
if args[0] < 0 then
return 0 - args[0]
end
return args[0]
end
let uf = ze_user_fn(name)
if uf != "" then
return ze_call_user(uf, args)
end
ze_rt_fail("unknown function " + name)
return false
end
make a function called ze_lookup takes env, name returns v
let i = to_num(list_len(env)) - 1
while i >= 0 do
if env[i][0] == name then
return env[i][1]
end
let i = i - 1
end
ze_rt_fail("unbound name " + name)
return false
end
make a function called ze_operands_ok takes op, a, b returns ok
let ta = type_of(a)
let tb = type_of(b)
if (ta == "list") or (tb == "list") or (ta == "bool") or (tb == "bool") then
ze_rt_fail("operator " + op + " needs numbers or strings, got " + ta + " and " + tb)
return false
end
if ta != tb then
ze_rt_fail("operator " + op + " needs operands of one type, got " + ta + " and " + tb)
return false
end
if (ta == "string") and ((op == "-") or (op == "*") or (op == "%")) then
ze_rt_fail("operator " + op + " is not defined on strings")
return false
end
return true
end
make a function called ze_eval takes ast, env returns v
let k = ast[0]
if k == "num" then
return ast[1]
end
if k == "str" then
return ast[1]
end
if k == "bool" then
return ast[1]
end
if k == "var" then
return ze_lookup(env, ast[1])
end
if k == "not" then
return ze_eval(ast[1], env) == false
end
if k == "neg" then
let v = ze_eval(ast[1], env)
if ze_operands_ok("-", 0, v) == false then
return false
end
return 0 - v
end
if k == "call" then
let args = []
let i = 0
let n = to_num(list_len(ast[2]))
while i < n do
let args = list_push(args, ze_eval(ast[2][i], env))
let i = i + 1
end
return ze_call(ast[1], args)
end
if (k == "forall") or (k == "exists") then
let dom = ze_eval(ast[2], env)
if type_of(dom) != "list" then
ze_rt_fail(k + " ranges over a set or a map's domain, got a " + type_of(dom))
return false
end
let i = 0
let n = to_num(list_len(dom))
while i < n do
let held = ze_eval(ast[3], list_appended(env, [ast[1], dom[i]]))
if (k == "forall") and (held == false) then
return false
end
if (k == "exists") and (held == true) then
return true
end
let i = i + 1
end
return k == "forall"
end
if k == "bin" then
let op = ast[1]
if op == "and" then
if ze_eval(ast[2], env) == false then
return false
end
return ze_eval(ast[3], env)
end
if op == "or" then
if ze_eval(ast[2], env) == true then
return true
end
return ze_eval(ast[3], env)
end
if op == "implies" then
if ze_eval(ast[2], env) == false then
return true
end
return ze_eval(ast[3], env)
end
let a = ze_eval(ast[2], env)
let b = ze_eval(ast[3], env)
if op == "==" then
return ze_val_eq(a, b)
end
if op == "!=" then
return ze_val_eq(a, b) == false
end
if op == "in" then
if type_of(b) != "list" then
ze_rt_fail("in needs a set on the right, got a " + type_of(b))
return false
end
return ze_elem_in(a, b)
end
if ze_operands_ok(op, a, b) == false then
return false
end
if op == "<" then
return a < b
end
if op == "<=" then
return a <= b
end
if op == ">" then
return a > b
end
if op == ">=" then
return a >= b
end
if op == "+" then
return a + b
end
if op == "-" then
return a - b
end
if op == "*" then
return a * b
end
if op == "%" then
return a % b
end
end
ze_rt_fail("cannot evaluate " + k)
return false
end
make a function called zs_registry returns registry
let existing = get("__vars", "zs_registry_obj")
if existing then
return existing
end
let registry = new("Dict", "zs_registry")
set_var("zs_registry_obj", registry)
return registry
end
make a function called zs_entry takes name returns entry
let e = get(zs_registry(), name)
if type_of(e) != "list" then
return ""
end
return e
end
make a function called zs_state_names takes name returns names
return zs_entry(name)[1]
end
make a function called zs_op_names takes name returns names
let ops = zs_entry(name)[6]
let out = []
let i = 0
let n = to_num(list_len(ops))
while i < n do
let out = list_push(out, ops[i][0])
let i = i + 1
end
return out
end
make a function called zs_find_op takes entry, op_name returns op
let ops = entry[6]
let i = 0
let n = to_num(list_len(ops))
while i < n do
if ops[i][0] == op_name then
return ops[i]
end
let i = i + 1
end
return ""
end
make a function called zs_is_query_op takes schema_name, op_name returns is_query
let entry = zs_entry(schema_name)
if type_of(entry) != "list" then
return false
end
let op = zs_find_op(entry, op_name)
if type_of(op) != "list" then
return false
end
return op[5] == true
end
make a function called zs_is_space takes c returns r
return (c == 32) or (c == 9) or (c == 13)
end
make a function called zs_trim takes s returns out
let h = str_intern(s)
let n = sc_len(h)
let a = 0
while (a < n) and zs_is_space(sc_code(h, a)) do
let a = a + 1
end
let b = n
while (b > a) and zs_is_space(sc_code(h, b - 1)) do
let b = b - 1
end
return sc_substr(h, a, b - a)
end
make a function called zs_lines takes text returns lines
let h = str_intern(text)
let n = sc_len(h)
let out = []
let buf = sb_new()
let lineno = 1
let i = 0
while i <= n do
let c = sc_code(h, i)
if (c == 10) or (c == -1) then
let raw = zs_trim(sb_str(buf))
let buf = sb_new()
if raw != "" then
let rh = str_intern(raw)
if sc_code(rh, 0) != 35 then
let rn = sc_len(rh)
let sp = 0
while (sp < rn) and (zs_is_space(sc_code(rh, sp)) == false) do
let sp = sp + 1
end
let kw = sc_substr(rh, 0, sp)
let rest = zs_trim(sc_substr(rh, sp, rn - sp))
let out = list_push(out, [lineno, kw, rest])
end
end
let lineno = lineno + 1
else
sb_push(buf, sc_char(h, i))
end
let i = i + 1
end
return out
end
make a function called zs_split_at takes text, sep returns parts
let h = str_intern(text)
let n = sc_len(h)
let idx = sm_find(h, n, sep, 0)
if idx < 0 then
return []
end
let sl = sc_len(str_intern(sep))
return [zs_trim(sc_substr(h, 0, idx)), zs_trim(sc_substr(h, idx + sl, n - idx - sl))]
end
make a function called zs_split_commas takes s returns names
let h = str_intern(s)
let n = sc_len(h)
let out = []
let buf = sb_new()
let i = 0
while i <= n do
let c = sc_code(h, i)
if (c == 44) or (c == -1) then
let piece = zs_trim(sb_str(buf))
let buf = sb_new()
if piece != "" then
let out = list_push(out, piece)
end
else
sb_push(buf, sc_char(h, i))
end
let i = i + 1
end
return out
end
make a function called zs_concat takes a, b returns out
let out = list_copy(a)
let i = 0
let n = to_num(list_len(b))
while i < n do
let out = list_push(out, b[i])
let i = i + 1
end
return out
end
make a function called zs_pair_names takes pairs returns names
let out = []
let i = 0
let n = to_num(list_len(pairs))
while i < n do
let out = list_push(out, pairs[i][0])
let i = i + 1
end
return out
end
make a function called zs_primed takes names returns out
let out = []
let i = 0
let n = to_num(list_len(names))
while i < n do
let out = list_push(out, names[i] + "'")
let i = i + 1
end
return out
end
make a function called zs_is_primed takes name returns r
let h = str_intern(name)
let n = sc_len(h)
if n == 0 then
return false
end
return sc_code(h, n - 1) == 39
end
make a function called zs_clause takes text, allowed returns result
let p = ze_parse(text)
if p[0] != "ok" then
return ["error", p[1]]
end
let names = ze_free_vars(p[1])
let i = 0
let n = to_num(list_len(names))
while i < n do
if ze_names_has(allowed, names[i]) == false then
if zs_is_primed(names[i]) then
return ["error", names[i] + " (after-state) is only meaningful in an ensure clause, or names no state variable"]
end
return ["error", "unbound name " + names[i] + " in: " + text]
end
let i = i + 1
end
return ["ok", p[1]]
end
make a function called zs_effect_var takes ast, state_names returns var
if ast[0] != "bin" then
return ""
end
if ast[1] != "==" then
return ""
end
if ast[2][0] != "var" then
return ""
end
let v = ast[2][1]
if zs_is_primed(v) == false then
return ""
end
let base = sc_substr(str_intern(v), 0, sc_len(str_intern(v)) - 1)
if ze_names_has(state_names, base) == false then
return ""
end
let fv = ze_free_vars(ast[3])
let i = 0
let n = to_num(list_len(fv))
while i < n do
if zs_is_primed(fv[i]) then
return ""
end
let i = i + 1
end
return base
end
make a function called zs_split_top_commas takes s returns parts
let h = str_intern(s)
let n = sc_len(h)
let out = []
let buf = sb_new()
let depth = 0
let quoted = false
let i = 0
while i < n do
let c = sc_code(h, i)
if c == 34 then
let quoted = quoted == false
end
if (quoted == false) and (c == 40) then
let depth = depth + 1
end
if (quoted == false) and (c == 41) then
let depth = depth - 1
end
if (quoted == false) and (depth == 0) and (c == 44) then
let out = list_push(out, zs_trim(sb_str(buf)))
let buf = sb_new()
else
sb_push(buf, sc_char(h, i))
end
let i = i + 1
end
let last = zs_trim(sb_str(buf))
if (last != "") or (to_num(list_len(out)) > 0) then
let out = list_push(out, last)
end
return out
end
make a function called zs_err takes lineno, msg returns result
return ["error", [lineno, msg]]
end
make a function called zs_eval_const takes tree, consts returns value
ze_rt_clear()
return ze_eval(tree, consts)
end
make a function called zs_load takes text returns result
let lines = zs_lines(text)
let n = to_num(list_len(lines))
let name = ""
let state = []
let consts = []
let inits = []
let invs = []
let domains = []
let ops = []
let facts = []
let goals = []
let negs = []
let acts = []
let dos = []
let thens = []
let refines_name = ""
let retrieves = []
let rejects = []
let in_op = false
let op_name = ""
let op_line = 0
let op_inputs = []
let op_pre = []
let op_post = []
let op_eff = []
let op_is_query = false
let i = 0
while i < n do
let lineno = lines[i][0]
let kw = lines[i][1]
let rest = lines[i][2]
if kw == "schema" then
if name != "" then
return zs_err(lineno, "only one schema per declaration")
end
if rest == "" then
return zs_err(lineno, "schema needs a name")
end
let name = rest
elif name == "" then
return zs_err(lineno, "a declaration starts with: schema Name")
elif kw == "state" then
let state = zs_concat(state, zs_split_commas(rest))
elif kw == "const" then
let parts = zs_split_at(rest, " = ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "const needs: const name = expr")
end
let c = zs_clause(parts[1], zs_pair_names(consts))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let v = zs_eval_const(c[1], consts)
if ze_rt_error() != "" then
return zs_err(lineno, ze_rt_error())
end
let consts = list_push(consts, [parts[0], v])
elif kw == "init" then
let parts = zs_split_at(rest, " = ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "init needs: init state_name = expr")
end
if ze_names_has(state, parts[0]) == false then
return zs_err(lineno, "init names " + parts[0] + ", which is not a declared state variable")
end
let c = zs_clause(parts[1], zs_pair_names(consts))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let v = zs_eval_const(c[1], consts)
if ze_rt_error() != "" then
return zs_err(lineno, ze_rt_error())
end
let inits = list_push(inits, [parts[0], v])
elif kw == "domain" then
let parts = zs_split_at(rest, " = ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "domain needs: domain input_name = expr")
end
let c = zs_clause(parts[1], zs_pair_names(consts))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let v = zs_eval_const(c[1], consts)
if ze_rt_error() != "" then
return zs_err(lineno, ze_rt_error())
end
if type_of(v) != "list" then
return zs_err(lineno, "a domain must be a set of values")
end
let domains = list_push(domains, [parts[0], v])
elif kw == "invariant" then
if in_op then
return zs_err(lineno, "invariant belongs to the schema, not inside an operation")
end
let c = zs_clause(rest, zs_concat(state, zs_pair_names(consts)))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let invs = list_push(invs, [rest, c[1]])
elif kw == "operation" then
if in_op then
return zs_err(lineno, "operation " + op_name + " is not closed with end")
end
let open_at = zs_split_at(rest, "(")
if to_num(list_len(open_at)) != 2 then
return zs_err(lineno, "operation needs: operation Name(input, ...)")
end
let close_at = zs_split_at(open_at[1], ")")
if to_num(list_len(close_at)) != 2 then
return zs_err(lineno, "operation is missing its closing )")
end
let in_op = true
let op_name = open_at[0]
let op_line = lineno
let op_inputs = zs_split_commas(close_at[0])
let op_pre = []
let op_post = []
let op_eff = []
let op_is_query = false
elif kw == "query" then
if in_op == false then
return zs_err(lineno, "query belongs inside an operation")
end
if (to_num(list_len(op_pre)) > 0) or (to_num(list_len(op_post)) > 0) then
return zs_err(lineno, "query must be the first line of an operation")
end
let op_is_query = true
elif kw == "require" then
if in_op == false then
return zs_err(lineno, "require belongs inside an operation")
end
let c = zs_clause(rest, zs_concat(zs_concat(state, op_inputs), zs_pair_names(consts)))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let op_pre = list_push(op_pre, [rest, c[1]])
elif kw == "ensure" then
if in_op == false then
return zs_err(lineno, "ensure belongs inside an operation")
end
let allowed = zs_concat(zs_concat(zs_concat(state, zs_primed(state)), op_inputs), zs_pair_names(consts))
let c = zs_clause(rest, allowed)
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let ev = zs_effect_var(c[1], state)
if ev != "" then
if op_is_query then
return zs_err(lineno, "query operation " + op_name + " may not define an after-state: " + rest)
end
let op_eff = list_push(op_eff, [ev, c[1][3], to_num(list_len(op_post))])
end
let op_post = list_push(op_post, [rest, c[1]])
elif kw == "end" then
if in_op == false then
return zs_err(lineno, "end without an open operation")
end
let ops = list_push(ops, [op_name, op_inputs, op_pre, op_post, op_eff, op_is_query])
let in_op = false
elif kw == "fact" then
let parts = zs_split_at(rest, " => ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "fact needs: fact template => expr")
end
let caps = sm_parse_template(parts[0])[1]
let c = zs_clause(parts[1], zs_concat(zs_concat(state, caps), zs_pair_names(consts)))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let facts = list_push(facts, [parts[0], parts[1], c[1]])
elif kw == "goal" then
let parts = zs_split_at(rest, " => ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "goal needs: goal text => OperationName")
end
let goals = list_push(goals, [parts[0], parts[1]])
elif kw == "negative" then
let negs = list_push(negs, rest)
elif kw == "action" then
let acts = list_push(acts, rest)
elif kw == "do" then
let parts = zs_split_at(rest, " => ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "do needs: do template => Operation(arg, ...)")
end
let caps = sm_parse_template(parts[0])[1]
let call = zs_split_at(parts[1], "(")
if to_num(list_len(call)) != 2 then
return zs_err(lineno, "do needs an operation call: Operation(arg, ...)")
end
let args_text = zs_trim(call[1])
let ah = str_intern(args_text)
let an = sc_len(ah)
if (an == 0) or (sc_code(ah, an - 1) != 41) then
return zs_err(lineno, "the operation call is missing its closing )")
end
let arg_texts = zs_split_top_commas(sc_substr(ah, 0, an - 1))
let arg_trees = []
let g = 0
let gn = to_num(list_len(arg_texts))
while g < gn do
let c = zs_clause(arg_texts[g], zs_concat(caps, zs_pair_names(consts)))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let arg_trees = list_push(arg_trees, c[1])
let g = g + 1
end
let dos = list_push(dos, [parts[0], call[0], arg_trees, arg_texts, lineno])
elif kw == "then" then
let parts = zs_split_at(rest, " => ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "then needs: then template => expr")
end
let caps = sm_parse_template(parts[0])[1]
let allowed = zs_concat(zs_concat(zs_concat(state, zs_primed(state)), caps), zs_pair_names(consts))
let c = zs_clause(parts[1], allowed)
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let thens = list_push(thens, [parts[0], parts[1], c[1]])
elif kw == "reject" then
let parts = zs_split_at(rest, " => ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "reject needs: reject template => the text of a require clause")
end
let rejects = list_push(rejects, [parts[0], parts[1], lineno])
elif kw == "refines" then
if rest == "" then
return zs_err(lineno, "refines needs the name of the schema it refines")
end
let refines_name = rest
elif kw == "retrieve" then
let parts = zs_split_at(rest, " = ")
if to_num(list_len(parts)) != 2 then
return zs_err(lineno, "retrieve needs: retrieve abstract_var = expr over this schema's state")
end
let c = zs_clause(parts[1], zs_concat(state, zs_pair_names(consts)))
if c[0] != "ok" then
return zs_err(lineno, c[1])
end
let retrieves = list_push(retrieves, [parts[0], parts[1], c[1]])
else
return zs_err(lineno, "unknown keyword " + kw)
end
let i = i + 1
end
if name == "" then
return zs_err(1, "a declaration starts with: schema Name")
end
if in_op then
return zs_err(op_line, "operation " + op_name + " is not closed with end")
end
let d = 0
let dn = to_num(list_len(dos))
while d < dn do
let found = false
let k = 0
let kn = to_num(list_len(ops))
while k < kn do
if ops[k][0] == dos[d][1] then
let found = true
if to_num(list_len(ops[k][1])) != to_num(list_len(dos[d][3])) then
return zs_err(dos[d][4], "do calls " + dos[d][1] + " with " + to_num(list_len(dos[d][3])) + " argument(s), and it takes " + to_num(list_len(ops[k][1])))
end
end
let k = k + 1
end
if found == false then
return zs_err(dos[d][4], "do names " + dos[d][1] + ", which is not an operation of " + name)
end
let d = d + 1
end
let rj = 0
let rjn = to_num(list_len(rejects))
while rj < rjn do
let found = false
let k = 0
let kn = to_num(list_len(ops))
while k < kn do
let q = 0
let qn = to_num(list_len(ops[k][2]))
while q < qn do
if ops[k][2][q][0] == rejects[rj][1] then
let found = true
end
let q = q + 1
end
let k = k + 1
end
if found == false then
return zs_err(rejects[rj][2], "reject names a clause no operation requires: " + rejects[rj][1])
end
let rj = rj + 1
end
let g = 0
let gn = to_num(list_len(goals))
while g < gn do
let known = false
let k = 0
let kn = to_num(list_len(ops))
while k < kn do
if ops[k][0] == goals[g][1] then
let known = true
end
let k = k + 1
end
if known == false then
return zs_err(1, "goal '" + goals[g][0] + "' names " + goals[g][1] + ", which is not an operation of " + name)
end
let g = g + 1
end
send(zs_registry(), "set", name, [name, state, consts, inits, invs, domains, ops, facts, goals, negs, acts, dos, thens, refines_name, retrieves, rejects])
return ["ok", name]
end
make a function called zs_env_add takes env, names, values, suffix returns out
let out = list_copy(env)
let i = 0
let n = to_num(list_len(names))
let m = to_num(list_len(values))
while (i < n) and (i < m) do
let out = list_push(out, [names[i] + suffix, values[i]])
let i = i + 1
end
return out
end
make a function called zs_first_failing takes clauses, env returns idx
let i = 0
let n = to_num(list_len(clauses))
while i < n do
if ze_eval(clauses[i][1], env) != true then
return i
end
let i = i + 1
end
return -1
end
make a function called zs_check_values takes schema_name, op_name, before, inputs, after returns diagnosis
let entry = zs_entry(schema_name)
if type_of(entry) != "list" then
return ["unknown_schema", [schema_name]]
end
let op = zs_find_op(entry, op_name)
if type_of(op) != "list" then
return ["unknown_operation", [schema_name, op_name]]
end
ze_rt_clear()
let state_names = entry[1]
let env_before = zs_env_add(entry[2], state_names, before, "")
let f = zs_first_failing(entry[4], env_before)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if f >= 0 then
return ["invariant_violated_before", [schema_name, entry[4][f][0], before]]
end
let env_op = zs_env_add(env_before, op[1], inputs, "")
let f = zs_first_failing(op[2], env_op)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if f >= 0 then
return ["precondition_violated", [op_name, op[2][f][0], before, inputs]]
end
if (to_num(list_len(state_names)) > 0) and (to_num(list_len(after)) == 0) then
return ["after_state_missing", [op_name]]
end
let env_after = zs_env_add(entry[2], state_names, after, "")
let f = zs_first_failing(entry[4], env_after)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if f >= 0 then
return ["invariant_violated_after", [schema_name, entry[4][f][0], after]]
end
let env_post = zs_env_add(env_op, state_names, after, "'")
let f = zs_first_failing(op[3], env_post)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if f >= 0 then
return ["postcondition_violated", [op_name, op[3][f][0], before, after, inputs]]
end
return ["ok", [op_name, before, after, inputs]]
end
make a function called zs_unbound takes v returns r
return (type_of(v) == "string") and (v == "")
end
make a function called zs_check takes schema_name, op_name returns diagnosis
let entry = zs_entry(schema_name)
if type_of(entry) != "list" then
return ["unknown_schema", [schema_name]]
end
let op = zs_find_op(entry, op_name)
if type_of(op) != "list" then
return ["unknown_operation", [schema_name, op_name]]
end
let names = entry[1]
let before = schema_state_values(schema_name, names)
let inputs = schema_input_values(schema_name, op_name, op[1])
let after = schema_state_values_after(schema_name, names)
let unbound_after = 0
let i = 0
let n = to_num(list_len(after))
while i < n do
if zs_unbound(after[i]) then
let unbound_after = unbound_after + 1
end
let i = i + 1
end
if (n > 0) and (unbound_after == n) then
let after = []
end
return zs_check_values(schema_name, op_name, before, inputs, after)
end
make a function called zs_format_diagnosis takes schema_name, op_name, diagnosis returns questions
let tag = diagnosis[0]
let p = diagnosis[1]
if tag == "invariant_violated_before" then
return ["The Given leaves " + schema_name + " in a state that breaks its invariant (" + p[1] + ") before " + op_name + " is even considered. Is the Given wrong, or is that invariant too strict?"]
end
if tag == "precondition_violated" then
return ["The scenario has " + op_name + " run from this Given, but its precondition clause (" + p[1] + ") is false. Is the Given wrong, is that clause too strict, or is the scenario about a rejection that needs its own treatment?"]
end
if tag == "invariant_violated_after" then
return ["The state the scenario claims after " + op_name + " breaks the invariant of " + schema_name + " (" + p[1] + "). Is the claimed outcome wrong, or does " + op_name + " need a stronger precondition to rule this case out?"]
end
if tag == "postcondition_violated" then
return ["The before and after states the scenario claims for " + op_name + " do not satisfy its postcondition (" + p[1] + "). Is the claimed outcome wrong, or is the postcondition?"]
end
if tag == "after_state_missing" then
return [op_name + " passed its precondition, so the scenario has to state the resulting state, and it stated none."]
end
if tag == "runtime_error" then
return ["A clause could not be evaluated: " + p[0]]
end
return []
end
make a function called zs_tick_ms returns ms
let v = get("__vars", "zs_tick_ms")
if type_of(v) == "unit" then
return 50
end
return v
end
make a function called zs_set_tick_ms takes ms returns done
set_var("zs_tick_ms", ms)
return true
end
make a function called zs_set_phase takes text returns done
set_var("zs_phase", text)
return true
end
make a function called zs_phase returns text
let v = get("__vars", "zs_phase")
if type_of(v) != "string" then
return ""
end
return v
end
make a function called zs_tick_poll takes tick_fn, a, b, c returns go
if tick_fn == "" then
return true
end
let since = get("__vars", "zs_tp_since")
if type_of(since) == "unit" then
let since = 0
end
let since = since + 1
if since < zs_tick_check_every() then
set_var("zs_tp_since", since)
return true
end
set_var("zs_tp_since", 0)
let now = now_ms()
let last = get("__vars", "zs_tp_last")
if type_of(last) == "unit" then
let last = 0
end
if now - last < zs_tick_ms() then
return true
end
set_var("zs_tp_last", now)
if apply(tick_fn, a, b, c) == false then
set_var("zs_aborted", 1)
return false
end
return true
end
make a function called zs_aborted returns r
return get("__vars", "zs_aborted") == 1
end
make a function called zs_clear_abort returns done
set_var("zs_aborted", 0)
return true
end
make a function called zs_tick_check_every returns n
return 32
end
make a function called zs_index_of takes names, name returns idx
let i = 0
let n = to_num(list_len(names))
while i < n do
if names[i] == name then
return i
end
let i = i + 1
end
return -1
end
make a function called zs_product takes lists returns tuples
let tuples = [[]]
let i = 0
let n = to_num(list_len(lists))
while i < n do
let nxt = []
let t = 0
let tn = to_num(list_len(tuples))
while t < tn do
let v = 0
let vn = to_num(list_len(lists[i]))
while v < vn do
let nxt = list_push(nxt, list_appended(tuples[t], lists[i][v]))
let v = v + 1
end
let t = t + 1
end
let tuples = nxt
let i = i + 1
end
return tuples
end
make a function called zs_domain_of takes entry, input returns values
let doms = entry[5]
let i = 0
let n = to_num(list_len(doms))
while i < n do
if doms[i][0] == input then
return doms[i][1]
end
let i = i + 1
end
return ""
end
make a function called zs_prepare takes entry returns result
return zs_prepare_mode(entry, true)
end
make a function called zs_prepare_mode takes entry, strict returns result
let state_names = entry[1]
let ns = to_num(list_len(state_names))
let ops = entry[6]
let prepared = []
let oi = 0
let opcount = to_num(list_len(ops))
while oi < opcount do
let op = ops[oi]
let doms = []
let k = 0
let kn = to_num(list_len(op[1]))
while (k < kn) and strict do
let d = zs_domain_of(entry, op[1][k])
if type_of(d) != "list" then
return ["missing_domain", [op[0], op[1][k]]]
end
let doms = list_push(doms, d)
let k = k + 1
end
let eff = []
let j = 0
while j < ns do
let eff = list_push(eff, "")
let j = j + 1
end
let used = []
let e = 0
let en = to_num(list_len(op[4]))
while e < en do
let idx = zs_index_of(state_names, op[4][e][0])
if eff[idx] == "" then
let eff = list_set(eff, idx, op[4][e][1])
let used = list_push(used, op[4][e][2])
end
let e = e + 1
end
let checked = []
let p = 0
let pn = to_num(list_len(op[3]))
while p < pn do
let is_effect = false
let u = 0
let un = to_num(list_len(used))
while u < un do
if used[u] == p then
let is_effect = true
end
let u = u + 1
end
if is_effect == false then
let fv = ze_free_vars(op[3][p][1])
let f = 0
let fcount = to_num(list_len(fv))
while f < fcount do
if zs_is_primed(fv[f]) then
let base = sc_substr(str_intern(fv[f]), 0, sc_len(str_intern(fv[f])) - 1)
let bi = zs_index_of(state_names, base)
if bi >= 0 then
if eff[bi] == "" then
return ["not_explorable", [op[0], base]]
end
end
end
let f = f + 1
end
let checked = list_push(checked, op[3][p])
end
let p = p + 1
end
let tuples = []
if strict then
let tuples = zs_product(doms)
end
let prepared = list_push(prepared, [op[0], op[1], tuples, op[2], checked, eff])
let oi = oi + 1
end
return ["ok", prepared]
end
make a function called zs_apply_effects takes eff, before, env_op returns after
let after = []
let j = 0
let n = to_num(list_len(before))
while j < n do
if eff[j] == "" then
let after = list_push(after, before[j])
else
let after = list_push(after, ze_eval(eff[j], env_op))
end
let j = j + 1
end
return after
end
make a function called zs_state_key takes state returns key
let b = sb_new()
let j = 0
let n = to_num(list_len(state))
while j < n do
sb_push(b, "" + j)
sb_push(b, ":")
sb_push(b, ze_canon(state[j]))
sb_push(b, ";")
let j = j + 1
end
return sb_str(b)
end
make a function called zs_trace takes states, parents, vias, idx returns steps
let rev = []
let cur = idx
while cur > 0 do
let rev = list_push(rev, [vias[cur][0], vias[cur][1], states[cur]])
let cur = parents[cur]
end
let steps = []
let i = to_num(list_len(rev)) - 1
while i >= 0 do
let steps = list_push(steps, rev[i])
let i = i - 1
end
return steps
end
make a function called zs_bfs takes schema_name, max_states, tick_fn returns result
let counter = get("__vars", "zs_bfs_counter")
if type_of(counter) == "unit" then
let counter = 0
end
set_var("zs_bfs_counter", counter + 1)
let seen = new("Dict", "zs_seen_" + counter)
let result = zs_bfs_impl(schema_name, max_states, tick_fn, seen)
object_delete(seen)
return result
end
make a function called zs_bfs_impl takes schema_name, max_states, tick_fn, seen returns result
let entry = zs_entry(schema_name)
if type_of(entry) != "list" then
return [["unknown_schema", [schema_name]], [], [], [], []]
end
let state_names = entry[1]
let ns = to_num(list_len(state_names))
let consts = entry[2]
let invariants = entry[4]
let init = []
let j = 0
while j < ns do
let found = false
let value = ""
let k = 0
let kn = to_num(list_len(entry[3]))
while k < kn do
if entry[3][k][0] == state_names[j] then
let found = true
let value = entry[3][k][1]
end
let k = k + 1
end
if found == false then
return [["missing_init", [state_names[j]]], [], [], [], []]
end
let init = list_push(init, value)
let j = j + 1
end
let prep = zs_prepare(entry)
if prep[0] != "ok" then
return [prep, [], [], [], []]
end
let ops = prep[1]
ze_rt_clear()
let f0 = zs_first_failing(invariants, zs_env_add(consts, state_names, init, ""))
if ze_rt_error() != "" then
return [["runtime_error", [ze_rt_error()]], [init], [-1], [[]], ops]
end
if f0 >= 0 then
return [["invariant_violated_init", [invariants[f0][0], init]], [init], [-1], [[]], ops]
end
let states = [init]
let parents = [-1]
let vias = [[]]
send(seen, "set", zs_state_key(init), 1)
let transitions = 0
let exhaustive = true
let counts = []
let cz = 0
while cz < to_num(list_len(ops)) do
let counts = list_push(counts, 0)
let cz = cz + 1
end
zs_clear_abort()
set_var("zs_tp_since", 0)
set_var("zs_tp_last", now_ms())
zs_set_phase("exploring the state space")
let qi = 0
let nops = to_num(list_len(ops))
while qi < to_num(list_len(states)) do
let st = states[qi]
let env_state = zs_env_add(consts, state_names, st, "")
let oi = 0
while oi < nops do
let op = ops[oi]
let tuples = op[2]
let ti = 0
let tn = to_num(list_len(tuples))
while ti < tn do
let inputs = tuples[ti]
let env_op = zs_env_add(env_state, op[1], inputs, "")
let fr = zs_first_failing(op[3], env_op)
if ze_rt_error() != "" then
return [["runtime_error", [ze_rt_error()]], states, parents, vias, ops]
end
if zs_tick_poll(tick_fn, qi, to_num(list_len(states)), transitions) == false then
return [["aborted", [to_num(list_len(states)), transitions, false]], states, parents, vias, ops]
end
if fr < 0 then
let transitions = transitions + 1
let counts = list_set(counts, oi, counts[oi] + 1)
let after = zs_apply_effects(op[5], st, env_op)
let fp = zs_first_failing(op[4], zs_env_add(env_op, state_names, after, "'"))
if ze_rt_error() != "" then
return [["runtime_error", [ze_rt_error()]], states, parents, vias, ops]
end
if fp >= 0 then
let trace = list_push(zs_trace(states, parents, vias, qi), [op[0], inputs, after])
return [["postcondition_violated", [op[0], op[4][fp][0], inputs, trace]], states, parents, vias, ops]
end
let fi = zs_first_failing(invariants, zs_env_add(consts, state_names, after, ""))
if ze_rt_error() != "" then
return [["runtime_error", [ze_rt_error()]], states, parents, vias, ops]
end
if fi >= 0 then
let trace = list_push(zs_trace(states, parents, vias, qi), [op[0], inputs, after])
return [["invariant_violated_after", [op[0], invariants[fi][0], inputs, trace]], states, parents, vias, ops]
end
let key = zs_state_key(after)
let known = get(seen, key)
if known != 1 then
if to_num(list_len(states)) >= max_states then
let exhaustive = false
else
send(seen, "set", key, 1)
let states = list_push(states, after)
let parents = list_push(parents, qi)
let vias = list_push(vias, [op[0], inputs])
end
end
end
let ti = ti + 1
end
let oi = oi + 1
end
let qi = qi + 1
end
let tag = "ok"
if exhaustive == false then
let tag = "ok_bounded"
end
return [[tag, [to_num(list_len(states)), transitions, exhaustive]], states, parents, vias, ops, counts]
end
make a function called zs_dead_ops takes schema_name, max_states returns names
let bfs = zs_bfs(schema_name, max_states, "")
if to_num(list_len(bfs)) < 6 then
return []
end
let out = []
let i = 0
let n = to_num(list_len(bfs[4]))
while i < n do
if bfs[5][i] == 0 then
let out = list_push(out, bfs[4][i][0])
end
let i = i + 1
end
return out
end
make a function called zs_explore_ticked takes schema_name, max_states, tick_fn returns diagnosis
return zs_bfs(schema_name, max_states, tick_fn)[0]
end
make a function called zs_explore takes schema_name, max_states returns diagnosis
return zs_bfs(schema_name, max_states, "")[0]
end
make a function called zs_format_exploration takes diagnosis returns lines
let tag = diagnosis[0]
let p = diagnosis[1]
if tag == "ok" then
return ["Explored every reachable state: " + p[0] + " states, " + p[1] + " transitions. No invariant or postcondition was violated."]
end
if tag == "ok_bounded" then
return ["Stopped at the bound after " + p[0] + " states (" + p[1] + " transitions) with no violation found. States beyond the bound were not checked."]
end
if tag == "invariant_violated_init" then
return ["The initial state already breaks an invariant: " + p[0]]
end
if (tag == "invariant_violated_after") or (tag == "postcondition_violated") then
let out = []
let what = "an invariant"
if tag == "postcondition_violated" then
let what = "a postcondition"
end
let out = list_push(out, p[0] + " reaches a state that breaks " + what + ": " + p[1])
let trace = p[3]
let i = 0
let n = to_num(list_len(trace))
while i < n do
let out = list_push(out, " step " + (i + 1) + ": " + trace[i][0] + " with " + ze_show(trace[i][1]))
let i = i + 1
end
return out
end
if tag == "missing_init" then
return ["The schema has no init for " + p[0] + ", so there is no state to start from."]
end
if tag == "missing_domain" then
return [p[0] + " takes " + p[1] + ", which has no domain declared, so its values cannot be enumerated."]
end
if tag == "not_explorable" then
return [p[0] + " constrains " + p[1] + "' without defining it (no ensure clause of the form " + p[1] + "' == expr), so its after-state is not determined."]
end
if tag == "aborted" then
return ["The search was stopped after " + p[0] + " states."]
end
if tag == "runtime_error" then
return ["A clause could not be evaluated: " + p[0]]
end
return ["Unrecognised result: " + tag]
end
make a function called zs_run_op takes entry, prep_op, state, inputs returns result
ze_rt_clear()
let env_state = zs_env_add(entry[2], entry[1], state, "")
let env_op = zs_env_add(env_state, prep_op[1], inputs, "")
let f = zs_first_failing(prep_op[3], env_op)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if f >= 0 then
return ["disabled", [prep_op[3][f][0]]]
end
let after = zs_apply_effects(prep_op[5], state, env_op)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
let fp = zs_first_failing(prep_op[4], zs_env_add(env_op, entry[1], after, "'"))
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if fp >= 0 then
return ["postcondition_violated", [prep_op[4][fp][0], after]]
end
let fi = zs_first_failing(entry[4], zs_env_add(entry[2], entry[1], after, ""))
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if fi >= 0 then
return ["invariant_violated_after", [entry[4][fi][0], after]]
end
return ["ok", [after]]
end
make a function called zs_find_prepared takes prep, op_name returns op
let i = 0
let n = to_num(list_len(prep))
while i < n do
if prep[i][0] == op_name then
return prep[i]
end
let i = i + 1
end
return ""
end
make a function called zs_apply takes schema_name, op_name, before, inputs returns result
let entry = zs_entry(schema_name)
if type_of(entry) != "list" then
return ["unknown_schema", [schema_name]]
end
let prep = zs_prepare_mode(entry, false)
if prep[0] != "ok" then
return prep
end
let op = zs_find_prepared(prep[1], op_name)
if type_of(op) != "list" then
return ["unknown_operation", [schema_name, op_name]]
end
return zs_run_op(entry, op, before, inputs)
end
make a function called zs_claim_check takes schema_name, op_name, before, inputs, claimed returns result
let r = zs_apply(schema_name, op_name, before, inputs)
if r[0] != "ok" then
return r
end
let after = r[1][0]
let names = zs_state_names(schema_name)
let i = 0
let n = to_num(list_len(names))
while (i < n) and (i < to_num(list_len(claimed))) do
if ze_val_eq(after[i], claimed[i]) == false then
return ["outcome_differs", [names[i], after[i], claimed[i]]]
end
let i = i + 1
end
return ["ok", [after]]
end
make a function called zs_failing_clauses takes clauses, env returns texts
let out = []
let i = 0
let n = to_num(list_len(clauses))
while i < n do
if ze_eval(clauses[i][1], env) != true then
let out = list_push(out, clauses[i][0])
end
let i = i + 1
end
return out
end
make a function called zf_parse_feature takes text returns feature
let lines = zs_lines(text)
let n = to_num(list_len(lines))
let goal_str = ""
let scenarios = []
let in_scenario = false
let name = ""
let steps = []
let phase = ""
let i = 0
while i < n do
let kw = lines[i][1]
let rest = lines[i][2]
let full = kw
if rest != "" then
let full = kw + " " + rest
end
if sc_code(str_intern(kw), 0) == 64 then
let i = i + 1
elif kw == "Feature:" then
let i = i + 1
elif (kw == "Scenario:") or (kw == "Scenario") then
if in_scenario then
let scenarios = list_push(scenarios, [name, steps])
end
let in_scenario = true
let name = rest
let steps = []
let phase = ""
let i = i + 1
elif (kw == "Given") or (kw == "When") or (kw == "Then") then
if kw == "Given" then
let phase = "given"
elif kw == "When" then
let phase = "when"
else
let phase = "then"
end
if in_scenario then
let steps = list_push(steps, [phase, rest])
end
let i = i + 1
elif (kw == "And") or (kw == "But") then
if in_scenario then
let steps = list_push(steps, [phase, rest])
end
let i = i + 1
else
if (in_scenario == false) and (goal_str == "") then
let h = str_intern(full)
let hn = sc_len(h)
let pos_at = sm_find(h, hn, " wants to ", 0)
if pos_at >= 0 then
let goal_str = zs_trim(sc_substr(h, pos_at + 10, hn - pos_at - 10))
else
if sm_at(h, hn, "I want to ", 0) then
let goal_str = zs_trim(sc_substr(h, 10, hn - 10))
end
end
end
let i = i + 1
end
end
if in_scenario then
let scenarios = list_push(scenarios, [name, steps])
end
return [goal_str, scenarios]
end
make a function called zf_literal takes s returns v
let h = str_intern(s)
let n = sc_len(h)
if n == 0 then
return s
end
let i = 0
if sc_code(h, 0) == 45 then
let i = 1
end
if i >= n then
return s
end
while i < n do
if ze_is_digit(sc_code(h, i)) == false then
return s
end
let i = i + 1
end
return to_num(s)
end
make a function called zf_match_fact takes entry, text returns hit
return zf_match_in(entry[7], text)
end
make a function called zf_match_in takes facts, text returns hit
let i = 0
let n = to_num(list_len(facts))
while i < n do
let m = sm_match(facts[i][0], text)
if m[0] == "ok" then
let caps = []
let c = 0
let cn = to_num(list_len(m[1]))
while c < cn do
let caps = list_push(caps, [m[1][c][0], zf_literal(m[1][c][1])])
let c = c + 1
end
return [i, caps]
end
let i = i + 1
end
return []
end
make a function called zf_matches_any takes templates, text returns r
let i = 0
let n = to_num(list_len(templates))
while i < n do
if sm_match(templates[i], text)[0] == "ok" then
return true
end
let i = i + 1
end
return false
end
make a function called zf_find_prepared takes prep, op_name returns op
let i = 0
let n = to_num(list_len(prep))
while i < n do
if prep[i][0] == op_name then
return prep[i]
end
let i = i + 1
end
return ""
end
make a function called zf_goal_op takes entry, goal_text returns op_name
let goals = entry[8]
let i = 0
let n = to_num(list_len(goals))
while i < n do
if goals[i][0] == goal_text then
return goals[i][1]
end
let i = i + 1
end
return ""
end
make a function called zf_non_input_caps takes caps, input_names returns out
let out = []
let i = 0
let n = to_num(list_len(caps))
while i < n do
if ze_names_has(input_names, caps[i][0]) == false then
let out = list_push(out, caps[i])
end
let i = i + 1
end
return out
end
make a function called zf_generalised takes template, input_names returns r
let caps = sm_parse_template(template)[1]
let i = 0
let n = to_num(list_len(caps))
while i < n do
if ze_names_has(input_names, caps[i]) == false then
return false
end
let i = i + 1
end
return true
end
make a function called zf_tuple_consistent takes tuple, input_names, caps returns r
let c = 0
let cn = to_num(list_len(caps))
while c < cn do
let pos_at = zs_index_of(input_names, caps[c][0])
if pos_at >= 0 then
if ze_val_eq(tuple[pos_at], caps[c][1]) == false then
return false
end
end
let c = c + 1
end
return true
end
make a function called zf_is_blocking takes tag returns r
if (tag == "unmapped_step") or (tag == "given_unreachable") or (tag == "contradicts_schema") or (tag == "requirement_not_in_schema") or (tag == "disjoint_from_schema") or (tag == "runtime_error") then
return true
end
if (tag == "outcome_contradicts_schema") or (tag == "rejection_contradicts_schema") or (tag == "rejection_reason_mismatch") or (tag == "operation_never_enabled") then
return true
end
if (tag == "sequence_step_disabled") or (tag == "sequence_step_violates") or (tag == "sequence_fact_fails") then
return true
end
return false
end
make a function called zf_any_blocking takes findings returns r
let i = 0
let n = to_num(list_len(findings))
while i < n do
if zf_is_blocking(findings[i][0]) then
return true
end
let i = i + 1
end
return false
end
make a function called zf_scenario_check takes entry, feature_steps, prep_op, states, parents, vias, bounded, tick_fn returns result
let findings = []
let matched = []
let i = 0
let n = to_num(list_len(feature_steps))
while i < n do
let phase = feature_steps[i][0]
let text = feature_steps[i][1]
if (phase == "given") or (phase == "when") then
let hit = zf_match_fact(entry, text)
if to_num(list_len(hit)) == 0 then
if zf_matches_any(entry[10], text) == false then
let findings = list_push(findings, ["unmapped_step", [text]])
end
else
let matched = list_push(matched, hit)
end
end
let i = i + 1
end
if zf_any_blocking(findings) then
return [findings, matched]
end
let consts = entry[2]
let state_names = entry[1]
let input_names = prep_op[1]
let all_caps = []
let m = 0
let mn = to_num(list_len(matched))
while m < mn do
let all_caps = zs_concat(all_caps, matched[m][1])
let m = m + 1
end
let sat = false
let witness = []
let si = 0
let sn = to_num(list_len(states))
while (si < sn) and (to_num(list_len(witness)) == 0) do
let env_state = zs_env_add(consts, state_names, states[si], "")
let tuples = prep_op[2]
let ti = 0
let tn = to_num(list_len(tuples))
while (ti < tn) and (to_num(list_len(witness)) == 0) do
if zf_tuple_consistent(tuples[ti], input_names, all_caps) then
if zs_tick_poll(tick_fn, si, sn, ti) == false then
return [findings, matched]
end
let env_op = zs_env_add(env_state, input_names, tuples[ti], "")
let env_facts = zs_concat(env_op, all_caps)
let holds = true
let f = 0
while (f < mn) and holds do
let ast = entry[7][matched[f][0]][2]
if ze_eval(ast, env_facts) != true then
let holds = false
end
let f = f + 1
end
if ze_rt_error() != "" then
return [list_push(findings, ["runtime_error", [ze_rt_error()]]), matched]
end
if holds then
let sat = true
if zs_first_failing(prep_op[3], env_op) < 0 then
let witness = [si, tuples[ti], zs_trace(states, parents, vias, si)]
end
end
end
let ti = ti + 1
end
let si = si + 1
end
if to_num(list_len(witness)) > 0 then
let findings = list_push(findings, ["witnessed", witness])
elif bounded then
let findings = list_push(findings, ["not_found_within_bound", []])
elif sat then
let findings = list_push(findings, ["contradicts_schema", []])
else
let findings = list_push(findings, ["given_unreachable", []])
end
return [findings, matched]
end
make a function called zf_necessity takes entry, fact_index, caps, prep_op, states, parents, vias, tick_fn returns finding
let consts = entry[2]
let state_names = entry[1]
let input_names = prep_op[1]
let fact = entry[7][fact_index]
let extra = zf_non_input_caps(caps, input_names)
let si = 0
let sn = to_num(list_len(states))
while si < sn do
let env_state = zs_env_add(consts, state_names, states[si], "")
let tuples = prep_op[2]
let ti = 0
let tn = to_num(list_len(tuples))
while ti < tn do
if zs_tick_poll(tick_fn, si, sn, ti) == false then
return []
end
let env_op = zs_env_add(env_state, input_names, tuples[ti], "")
if zs_first_failing(prep_op[3], env_op) < 0 then
if ze_eval(fact[2], zs_concat(env_op, extra)) != true then
return ["requirement_not_in_schema", [fact[0], fact[1], zs_trace(states, parents, vias, si), tuples[ti]]]
end
end
let ti = ti + 1
end
let si = si + 1
end
return []
end
make a function called zf_scenario_mode takes entry, steps returns mode
let neg = false
let seq = false
let i = 0
let n = to_num(list_len(steps))
while i < n do
let phase = steps[i][0]
let text = steps[i][1]
if zf_matches_any(entry[9], text) then
let neg = true
end
if to_num(list_len(zf_match_in(entry[15], text))) > 0 then
let neg = true
end
if phase != "then" then
if to_num(list_len(zf_match_in(entry[11], text))) > 0 then
let seq = true
end
end
let i = i + 1
end
if neg then
return "negative"
end
if seq then
return "sequence"
end
return "positive"
end
make a function called zf_reject_clause takes entry, steps returns clause
let i = 0
let n = to_num(list_len(steps))
while i < n do
let hit = zf_match_in(entry[15], steps[i][1])
if to_num(list_len(hit)) > 0 then
return entry[15][hit[0]][1]
end
let i = i + 1
end
return ""
end
make a function called zf_collect_thens takes entry, steps returns result
let hits = []
let findings = []
let i = 0
let n = to_num(list_len(steps))
while i < n do
if steps[i][0] == "then" then
let hit = zf_match_in(entry[12], steps[i][1])
if to_num(list_len(hit)) == 0 then
let findings = list_push(findings, ["unchecked_outcome", [steps[i][1]]])
else
let hits = list_push(hits, hit)
end
end
let i = i + 1
end
return [hits, findings]
end
make a function called zf_outcome_check takes entry, matched, then_hits, prep_op, states, parents, vias, tick_fn returns findings
if to_num(list_len(then_hits)) == 0 then
return []
end
let consts = entry[2]
let state_names = entry[1]
let input_names = prep_op[1]
let all_caps = []
let m = 0
let mn = to_num(list_len(matched))
while m < mn do
let all_caps = zs_concat(all_caps, matched[m][1])
let m = m + 1
end
let q = 0
let qn = to_num(list_len(then_hits))
while q < qn do
let all_caps = zs_concat(all_caps, then_hits[q][1])
let q = q + 1
end
ze_rt_clear()
let count = 0
let si = 0
let sn = to_num(list_len(states))
while si < sn do
let env_state = zs_env_add(consts, state_names, states[si], "")
let tuples = prep_op[2]
let ti = 0
let tn = to_num(list_len(tuples))
while ti < tn do
if zf_tuple_consistent(tuples[ti], input_names, all_caps) then
if zs_tick_poll(tick_fn, si, sn, ti) == false then
return []
end
let env_op = zs_env_add(env_state, input_names, tuples[ti], "")
let env_facts = zs_concat(env_op, all_caps)
let holds = true
let f = 0
while (f < mn) and holds do
if ze_eval(entry[7][matched[f][0]][2], env_facts) != true then
let holds = false
end
let f = f + 1
end
if holds and (zs_first_failing(prep_op[3], env_op) < 0) then
let r = zs_run_op(entry, prep_op, states[si], tuples[ti])
if r[0] == "ok" then
let after = r[1][0]
let env_then = zs_env_add(env_facts, state_names, after, "'")
let t = 0
while t < qn do
let tf = entry[12][then_hits[t][0]]
if ze_eval(tf[2], env_then) != true then
if ze_rt_error() != "" then
return [["runtime_error", [ze_rt_error()]]]
end
return [["outcome_contradicts_schema", [tf[0], tf[1], zs_trace(states, parents, vias, si), tuples[ti], after]]]
end
let t = t + 1
end
let count = count + 1
end
end
end
let ti = ti + 1
end
let si = si + 1
end
if ze_rt_error() != "" then
return [["runtime_error", [ze_rt_error()]]]
end
if count == 0 then
return []
end
return [["outcome_confirmed", [count]]]
end
make a function called zf_negative_check takes entry, steps, prep_op, states, parents, vias, bounded, tick_fn returns findings
let findings = []
let matched = []
let i = 0
let n = to_num(list_len(steps))
while i < n do
if steps[i][0] != "then" then
let hit = zf_match_in(entry[7], steps[i][1])
if to_num(list_len(hit)) == 0 then
if zf_matches_any(entry[10], steps[i][1]) == false then
let findings = list_push(findings, ["unmapped_step", [steps[i][1]]])
end
else
let matched = list_push(matched, hit)
end
end
let i = i + 1
end
if zf_any_blocking(findings) then
return findings
end
let reason = zf_reject_clause(entry, steps)
let consts = entry[2]
let state_names = entry[1]
let input_names = prep_op[1]
let all_caps = []
let m = 0
let mn = to_num(list_len(matched))
while m < mn do
let all_caps = zs_concat(all_caps, matched[m][1])
let m = m + 1
end
ze_rt_clear()
let sat = false
let disabled_count = 0
let accepted_count = 0
let witness = []
let reason_hit = false
let sample = []
let si = 0
let sn = to_num(list_len(states))
while si < sn do
let env_state = zs_env_add(consts, state_names, states[si], "")
let tuples = prep_op[2]
let ti = 0
let tn = to_num(list_len(tuples))
while ti < tn do
if zf_tuple_consistent(tuples[ti], input_names, all_caps) then
if zs_tick_poll(tick_fn, si, sn, ti) == false then
return findings
end
let env_op = zs_env_add(env_state, input_names, tuples[ti], "")
let env_facts = zs_concat(env_op, all_caps)
let holds = true
let f = 0
while (f < mn) and holds do
if ze_eval(entry[7][matched[f][0]][2], env_facts) != true then
let holds = false
end
let f = f + 1
end
if holds then
let sat = true
let failing = zs_failing_clauses(prep_op[3], env_op)
if to_num(list_len(failing)) == 0 then
let accepted_count = accepted_count + 1
else
let disabled_count = disabled_count + 1
if to_num(list_len(witness)) == 0 then
let witness = [si, tuples[ti], failing, zs_trace(states, parents, vias, si)]
end
if reason != "" then
if ze_names_has(failing, reason) then
let reason_hit = true
end
if to_num(list_len(sample)) == 0 then
let sample = failing
end
end
end
end
end
let ti = ti + 1
end
let si = si + 1
end
if ze_rt_error() != "" then
return list_appended(findings, ["runtime_error", [ze_rt_error()]])
end
if sat == false then
if bounded then
return list_appended(findings, ["not_found_within_bound", []])
end
return list_appended(findings, ["given_unreachable", []])
end
if disabled_count == 0 then
if bounded then
return list_appended(findings, ["not_found_within_bound", []])
end
return list_appended(findings, ["rejection_contradicts_schema", [accepted_count]])
end
let findings = list_appended(findings, ["rejection_witnessed", [witness[0], witness[1], witness[2], witness[3], disabled_count, accepted_count]])
if reason != "" then
if reason_hit then
let findings = list_appended(findings, ["rejection_reason_confirmed", [reason]])
else
let findings = list_appended(findings, ["rejection_reason_mismatch", [reason, sample]])
end
end
return findings
end
make a function called zf_sequence_check takes entry, steps, prep, states, parents, vias, bounded, tick_fn returns findings
let findings = []
let items = []
let i = 0
let n = to_num(list_len(steps))
while i < n do
let phase = steps[i][0]
let text = steps[i][1]
if phase == "then" then
let hit = zf_match_in(entry[12], text)
if to_num(list_len(hit)) == 0 then
let findings = list_push(findings, ["unchecked_outcome", [text]])
else
let items = list_push(items, ["then", hit[0], hit[1], text])
end
else
let dohit = []
if phase == "when" then
let dohit = zf_match_in(entry[11], text)
end
if to_num(list_len(dohit)) > 0 then
let items = list_push(items, ["do", dohit[0], dohit[1], text])
else
let fh = zf_match_in(entry[7], text)
if to_num(list_len(fh)) == 0 then
if zf_matches_any(entry[10], text) == false then
let findings = list_push(findings, ["unmapped_step", [text]])
end
else
let items = list_push(items, ["fact", fh[0], fh[1], text])
end
end
end
let i = i + 1
end
if zf_any_blocking(findings) then
return findings
end
let consts = entry[2]
let state_names = entry[1]
let frontier = []
let si = 0
let sn = to_num(list_len(states))
while si < sn do
let frontier = list_push(frontier, [states[si], states[si], si, []])
let si = si + 1
end
let ran = false
let k = 0
let kn = to_num(list_len(items))
while k < kn do
let it = items[k]
let kind = it[0]
let text = it[3]
let caps = it[2]
let next_frontier = []
ze_rt_clear()
if kind == "fact" then
let ast = entry[7][it[1]][2]
let e = 0
let en = to_num(list_len(frontier))
while e < en do
if zs_tick_poll(tick_fn, e, en, k) == false then
return findings
end
let env = zs_concat(zs_env_add(consts, state_names, frontier[e][0], ""), caps)
if ze_eval(ast, env) == true then
let next_frontier = list_push(next_frontier, frontier[e])
end
let e = e + 1
end
if ze_rt_error() != "" then
return list_appended(findings, ["runtime_error", [ze_rt_error()]])
end
if to_num(list_len(next_frontier)) == 0 then
if bounded then
return list_appended(findings, ["not_found_within_bound", [text]])
end
if ran then
return list_appended(findings, ["sequence_fact_fails", [text]])
end
return list_appended(findings, ["given_unreachable", []])
end
elif kind == "do" then
let dentry = entry[11][it[1]]
let op = zs_find_prepared(prep, dentry[1])
let env_c = zs_concat(consts, caps)
let args = []
let g = 0
let gn = to_num(list_len(dentry[2]))
while g < gn do
let args = list_push(args, ze_eval(dentry[2][g], env_c))
let g = g + 1
end
if ze_rt_error() != "" then
return list_appended(findings, ["runtime_error", [ze_rt_error()]])
end
let sample_clause = ""
let sample_state = []
let violation = []
let e = 0
let en = to_num(list_len(frontier))
while e < en do
if zs_tick_poll(tick_fn, e, en, k) == false then
return findings
end
let r = zs_run_op(entry, op, frontier[e][0], args)
if r[0] == "ok" then
let next_frontier = list_push(next_frontier, [r[1][0], frontier[e][0], frontier[e][2], list_appended(frontier[e][3], [dentry[1], args])])
elif r[0] == "disabled" then
if sample_clause == "" then
let sample_clause = r[1][0]
let sample_state = frontier[e][0]
end
else
if to_num(list_len(violation)) == 0 then
let violation = [r[0], r[1]]
end
end
let e = e + 1
end
if to_num(list_len(next_frontier)) == 0 then
if bounded then
return list_appended(findings, ["not_found_within_bound", [text]])
end
if sample_clause != "" then
return list_appended(findings, ["sequence_step_disabled", [text, sample_clause, sample_state]])
end
return list_appended(findings, ["sequence_step_violates", [text, violation[0], violation[1]]])
end
let ran = true
else
let tf = entry[12][it[1]]
let e = 0
let en = to_num(list_len(frontier))
while e < en do
if zs_tick_poll(tick_fn, e, en, k) == false then
return findings
end
let env_before = zs_env_add(consts, state_names, frontier[e][1], "")
let env = zs_concat(zs_env_add(env_before, state_names, frontier[e][0], "'"), caps)
if ze_eval(tf[2], env) == true then
let next_frontier = list_push(next_frontier, frontier[e])
end
let e = e + 1
end
if ze_rt_error() != "" then
return list_appended(findings, ["runtime_error", [ze_rt_error()]])
end
if to_num(list_len(next_frontier)) == 0 then
if bounded then
return list_appended(findings, ["not_found_within_bound", [text]])
end
return list_appended(findings, ["outcome_contradicts_schema", [tf[0], tf[1], frontier[0][1], frontier[0][0]]])
end
end
let frontier = next_frontier
let k = k + 1
end
let first = frontier[0]
return list_appended(findings, ["sequence_realised", [to_num(list_len(frontier)), first[2], zs_trace(states, parents, vias, first[2]), first[3]]])
end
make a function called zs_analyse_feature takes schema_name, feature_text, max_states returns result
return zs_analyse_feature_ticked(schema_name, feature_text, max_states, "")
end
make a function called zs_analyse_feature_ticked takes schema_name, feature_text, max_states, tick_fn returns result
let entry = zs_entry(schema_name)
if type_of(entry) != "list" then
return ["unknown_schema", [schema_name]]
end
let feature = zf_parse_feature(feature_text)
let goal_text = feature[0]
let op_name = zf_goal_op(entry, goal_text)
let scenarios = feature[1]
let needs_goal = false
let mi = 0
let mn = to_num(list_len(scenarios))
while mi < mn do
if zf_scenario_mode(entry, scenarios[mi][1]) != "sequence" then
let needs_goal = true
end
let mi = mi + 1
end
if needs_goal and (op_name == "") then
return ["unmapped_goal", [goal_text]]
end
let bfs = zs_bfs(schema_name, max_states, tick_fn)
let diag = bfs[0]
if diag[0] == "aborted" then
return ["aborted", ["exploring the state space"]]
end
if (diag[0] != "ok") and (diag[0] != "ok_bounded") then
return ["schema_unsound", [diag]]
end
let bounded = diag[0] == "ok_bounded"
let states = bfs[1]
let parents = bfs[2]
let vias = bfs[3]
let prep_op = ""
let input_names = []
if op_name != "" then
let prep_op = zf_find_prepared(bfs[4], op_name)
let input_names = prep_op[1]
end
let goal_findings = []
if bounded then
let goal_findings = list_push(goal_findings, ["bound_reached", [to_num(list_len(states))]])
end
if (op_name != "") and (bounded == false) then
let oi = 0
let on_total = to_num(list_len(bfs[4]))
while oi < on_total do
if (bfs[4][oi][0] == op_name) and (bfs[5][oi] == 0) then
let goal_findings = list_push(goal_findings, ["operation_never_enabled", [op_name]])
end
let oi = oi + 1
end
end
let reports = []
let obligations = []
let seen_keys = []
let si = 0
let sn = to_num(list_len(scenarios))
while si < sn do
let name = scenarios[si][0]
let steps = scenarios[si][1]
let mode = zf_scenario_mode(entry, steps)
zs_set_phase("checking scenario: " + name)
if mode == "negative" then
let fs = zf_negative_check(entry, steps, prep_op, states, parents, vias, bounded, tick_fn)
if zs_aborted() then
return ["aborted", [zs_phase()]]
end
let reports = list_push(reports, [name, "negative", fs])
elif mode == "sequence" then
let fs = zf_sequence_check(entry, steps, bfs[4], states, parents, vias, bounded, tick_fn)
if zs_aborted() then
return ["aborted", [zs_phase()]]
end
let reports = list_push(reports, [name, "sequence", fs])
else
let r = zf_scenario_check(entry, steps, prep_op, states, parents, vias, bounded, tick_fn)
if zs_aborted() then
return ["aborted", [zs_phase()]]
end
let fs = r[0]
let thens = zf_collect_thens(entry, steps)
let fs = zs_concat(fs, thens[1])
if zf_any_blocking(r[0]) == false then
let fs = zs_concat(fs, zf_outcome_check(entry, r[1], thens[0], prep_op, states, parents, vias, tick_fn))
if zs_aborted() then
return ["aborted", [zs_phase()]]
end
end
let reports = list_push(reports, [name, "positive", fs])
let m = 0
let mcount = to_num(list_len(r[1]))
while m < mcount do
let idx = r[1][m][0]
let caps = r[1][m][1]
let template = entry[7][idx][0]
let key = template
if zf_generalised(template, input_names) == false then
let key = template + " @ " + name
end
if ze_names_has(seen_keys, key) == false then
let seen_keys = list_push(seen_keys, key)
let obligations = list_push(obligations, [idx, caps])
end
let m = m + 1
end
end
let si = si + 1
end
let o = 0
let on_count = to_num(list_len(obligations))
while o < on_count do
let idx = obligations[o][0]
let fact = entry[7][idx]
let vocab = zs_concat(zs_concat(entry[1], zs_pair_names(entry[2])), input_names)
let fv = ze_free_vars(fact[2])
let touches = false
let v = 0
let vn = to_num(list_len(fv))
while v < vn do
if ze_names_has(vocab, fv[v]) then
let touches = true
end
let v = v + 1
end
if touches == false then
let goal_findings = list_push(goal_findings, ["disjoint_from_schema", [fact[0]]])
end
ze_rt_clear()
zs_set_phase("checking that " + zf_fact_name(entry, idx) + " is enforced")
let nf = zf_necessity(entry, idx, obligations[o][1], prep_op, states, parents, vias, tick_fn)
if zs_aborted() then
return ["aborted", [zs_phase()]]
end
if ze_rt_error() != "" then
let goal_findings = list_push(goal_findings, ["runtime_error", [ze_rt_error()]])
elif to_num(list_len(nf)) > 0 then
let goal_findings = list_push(goal_findings, nf)
end
let o = o + 1
end
let blocking = zf_any_blocking(goal_findings)
let ri = 0
let rn = to_num(list_len(reports))
while ri < rn do
if zf_any_blocking(reports[ri][2]) then
let blocking = true
end
let ri = ri + 1
end
let tag = "consistent"
if blocking then
let tag = "inconsistent"
end
let stats = [to_num(list_len(states)), diag[1][1], bounded == false]
return [tag, [goal_text, op_name, reports, goal_findings, stats]]
end
make a function called zf_fact_name takes entry, idx returns text
return "'" + entry[7][idx][0] + "'"
end
make a function called zf_describe_finding takes finding returns lines
let tag = finding[0]
let p = finding[1]
if tag == "unmapped_step" then
return [" no fact is declared for the step: " + p[0]]
end
if tag == "given_unreachable" then
return [" no reachable state satisfies this scenario's Given and When facts together"]
end
if tag == "contradicts_schema" then
return [" states satisfy the facts, but the operation is never enabled in any of them: the scenario claims what the schema does not allow"]
end
if tag == "not_found_within_bound" then
return [" no witness found, but the search stopped at its bound, so this is not conclusive"]
end
if tag == "witnessed" then
let out = [" realisable: reached by " + to_num(list_len(p[2])) + " step(s), then the operation runs with " + ze_show(p[1])]
return out
end
if tag == "negative_skipped" then
return [" negative scenario, not checked (the thesis excludes negative scenarios)"]
end
if tag == "outcome_confirmed" then
return [" outcome confirmed: the schema computes the claimed result in all " + p[0] + " instance(s) the scenario covers"]
end
if tag == "outcome_contradicts_schema" then
return [" the schema computes a different outcome from the claim: \"" + p[0] + "\" (" + p[1] + ")"]
end
if tag == "unchecked_outcome" then
return [" outcome not checked, no meaning is declared for: " + p[0]]
end
if tag == "rejection_witnessed" then
return [" rejection supported: the operation is disabled in " + p[4] + " instance(s) and accepted in " + p[5] + "; failing clause(s) in the first: " + zsr_join(p[2])]
end
if tag == "rejection_reason_confirmed" then
return [" the named reason is among the failing clauses: " + p[0]]
end
if tag == "rejection_reason_mismatch" then
return [" the named reason (" + p[0] + ") never fails where the operation is disabled; in the first disabled instance the failing clause(s) are: " + zsr_join(p[1])]
end
if tag == "rejection_contradicts_schema" then
return [" the schema accepts the operation in all " + p[0] + " instance(s) the facts cover, so the scenario's rejection never happens"]
end
if tag == "sequence_realised" then
return [" realised by running the schema from " + p[0] + " starting state(s); one run: " + zsr_ops(p[3])]
end
if tag == "sequence_step_disabled" then
return [" the step \"" + p[0] + "\" is not enabled; the clause that fails is: " + p[1]]
end
if tag == "sequence_step_violates" then
return [" the step \"" + p[0] + "\" runs but breaks a clause of the schema (" + p[1] + ")"]
end
if tag == "sequence_fact_fails" then
return [" no state reached by the earlier steps satisfies: " + p[0]]
end
return [" " + tag]
end
make a function called zsr_join takes items returns text
let b = sb_new()
let i = 0
let n = to_num(list_len(items))
while i < n do
if i > 0 then
sb_push(b, "; ")
end
sb_push(b, items[i])
let i = i + 1
end
return sb_str(b)
end
make a function called zsr_ops takes ran returns text
let b = sb_new()
let i = 0
let n = to_num(list_len(ran))
while i < n do
if i > 0 then
sb_push(b, " then ")
end
sb_push(b, ran[i][0])
sb_push(b, " with ")
sb_push(b, ze_show(ran[i][1]))
let i = i + 1
end
return sb_str(b)
end
make a function called zs_format_analysis takes result returns lines
let tag = result[0]
let p = result[1]
if tag == "unknown_schema" then
return ["No schema named " + p[0] + " is loaded."]
end
if tag == "aborted" then
return ["The analysis was stopped while " + p[0] + "."]
end
if tag == "unmapped_goal" then
return ["The feature's goal, \"" + p[0] + "\", is not mapped to an operation (add: goal " + p[0] + " => OperationName)."]
end
if tag == "schema_unsound" then
let out = ["The schema breaks its own invariants, so no feature can be checked against it."]
return zs_concat(out, zs_format_exploration(p[0]))
end
let out = []
let out = list_push(out, "Goal: " + p[0] + " (goal-related operation: " + p[1] + ")")
let out = list_push(out, "Schema states explored: " + p[4][0] + ", exhaustive: " + p[4][2])
let reports = p[2]
let i = 0
let n = to_num(list_len(reports))
while i < n do
let out = list_push(out, "Scenario (" + reports[i][1] + "): " + reports[i][0])
let f = 0
let fn_count = to_num(list_len(reports[i][2]))
while f < fn_count do
let out = zs_concat(out, zf_describe_finding(reports[i][2][f]))
let f = f + 1
end
let i = i + 1
end
let gf = p[3]
let g = 0
let gn = to_num(list_len(gf))
while g < gn do
let fd = gf[g]
if fd[0] == "requirement_not_in_schema" then
let out = list_push(out, "Requirement missing from the schema: \"" + fd[1][0] + "\" (" + fd[1][1] + ")")
let out = list_push(out, " " + p[1] + " is enabled with inputs " + ze_show(fd[1][3]) + " in a state reached by " + to_num(list_len(fd[1][2])) + " step(s), where that fact is false")
elif fd[0] == "operation_never_enabled" then
let out = list_push(out, "The goal operation " + fd[1][0] + " is not enabled in any state the schema can reach.")
elif fd[0] == "disjoint_from_schema" then
let out = list_push(out, "Fact reads nothing of the schema: \"" + fd[1][0] + "\"")
elif fd[0] == "bound_reached" then
let out = list_push(out, "The search stopped at its bound after " + fd[1][0] + " states; the checks cover only those states.")
else
let out = list_push(out, "Finding: " + fd[0])
end
let g = g + 1
end
if tag == "consistent" then
let out = list_push(out, "Consistent: every fact is enforced by " + p[1] + ", and each scenario holds when the schema runs it.")
else
let out = list_push(out, "Inconsistent: the feature and the schema disagree.")
end
return out
end
make a function called zr_find_retrieve takes retrieves, var returns tree
let i = 0
let n = to_num(list_len(retrieves))
while i < n do
if retrieves[i][0] == var then
return retrieves[i][2]
end
let i = i + 1
end
return ""
end
make a function called zr_retrieved takes retrieves, abstract_names, env returns values
let out = []
let i = 0
let n = to_num(list_len(abstract_names))
while i < n do
let out = list_push(out, ze_eval(zr_find_retrieve(retrieves, abstract_names[i]), env))
let i = i + 1
end
return out
end
make a function called zr_find_op takes ops, name returns op
let i = 0
let n = to_num(list_len(ops))
while i < n do
if ops[i][0] == name then
return ops[i]
end
let i = i + 1
end
return ""
end
make a function called zr_shape takes concrete_name returns result
let ec = zs_entry(concrete_name)
if type_of(ec) != "list" then
return ["unknown_schema", [concrete_name]]
end
if ec[13] == "" then
return ["not_a_refinement", [concrete_name]]
end
let ea = zs_entry(ec[13])
if type_of(ea) != "list" then
return ["unknown_abstract", [ec[13]]]
end
let i = 0
let n = to_num(list_len(ea[1]))
while i < n do
if type_of(zr_find_retrieve(ec[14], ea[1][i])) != "list" then
return ["retrieve_incomplete", [ea[1][i]]]
end
let i = i + 1
end
let j = 0
let jn = to_num(list_len(ea[6]))
while j < jn do
let cop = zr_find_op(ec[6], ea[6][j][0])
if type_of(cop) != "list" then
return ["operation_missing", [ea[6][j][0]]]
end
if ze_val_eq(cop[1], ea[6][j][1]) == false then
return ["input_mismatch", [ea[6][j][0], ea[6][j][1], cop[1]]]
end
let j = j + 1
end
return ["ok", ea]
end
make a function called zs_refines takes concrete_name, max_states, tick_fn returns result
let shape = zr_shape(concrete_name)
if shape[0] != "ok" then
return shape
end
let ec = zs_entry(concrete_name)
let ea = shape[1]
let anames = ea[1]
let ainit = []
let i = 0
let n = to_num(list_len(anames))
while i < n do
let found = false
let value = ""
let k = 0
let kn = to_num(list_len(ea[3]))
while k < kn do
if ea[3][k][0] == anames[i] then
let found = true
let value = ea[3][k][1]
end
let k = k + 1
end
if found == false then
return ["abstract_missing_init", [anames[i]]]
end
let ainit = list_push(ainit, value)
let i = i + 1
end
let prep_a = zs_prepare_mode(ea, false)
if prep_a[0] != "ok" then
return prep_a
end
let bfs = zs_bfs(concrete_name, max_states, tick_fn)
let diag = bfs[0]
if diag[0] == "aborted" then
return ["aborted", [diag[1][0]]]
end
if (diag[0] != "ok") and (diag[0] != "ok_bounded") then
return ["concrete_unsound", [diag]]
end
let states = bfs[1]
let parents = bfs[2]
let vias = bfs[3]
let prep_c = bfs[4]
let cconsts = ec[2]
let cnames = ec[1]
let aconsts = ea[2]
zs_set_phase("checking the refinement")
ze_rt_clear()
let checks = 0
let ci = 0
let cn = to_num(list_len(states))
while ci < cn do
let c = states[ci]
let env_c = zs_env_add(cconsts, cnames, c, "")
let a = zr_retrieved(ec[14], anames, env_c)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if ci == 0 then
let v = 0
while v < n do
if ze_val_eq(a[v], ainit[v]) == false then
return ["init_not_refined", [anames[v], a[v], ainit[v]]]
end
let v = v + 1
end
end
let env_a = zs_env_add(aconsts, anames, a, "")
let fi = zs_first_failing(ea[4], env_a)
if fi >= 0 then
return ["retrieve_breaks_invariant", [ea[4][fi][0], zs_trace(states, parents, vias, ci)]]
end
let oi = 0
let on_total = to_num(list_len(prep_c))
while oi < on_total do
let pc = prep_c[oi]
let pa = zs_find_prepared(prep_a[1], pc[0])
let tuples = pc[2]
let ti = 0
let tn = to_num(list_len(tuples))
while ti < tn do
if zs_tick_poll(tick_fn, ci, cn, ti) == false then
return ["aborted", [cn]]
end
let env_op_c = zs_env_add(env_c, pc[1], tuples[ti], "")
let fc = zs_first_failing(pc[3], env_op_c)
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
if type_of(pa) == "list" then
let env_op_a = zs_env_add(env_a, pa[1], tuples[ti], "")
let fa = zs_first_failing(pa[3], env_op_a)
if (fa < 0) and (fc >= 0) then
return ["not_applicable", [pc[0], tuples[ti], pc[3][fc][0], zs_trace(states, parents, vias, ci)]]
end
if (fa < 0) and (fc < 0) then
let after_c = zs_apply_effects(pc[5], c, env_op_c)
let after_a = zs_apply_effects(pa[5], a, env_op_a)
let back = zr_retrieved(ec[14], anames, zs_env_add(cconsts, cnames, after_c, ""))
let v = 0
while v < n do
if ze_val_eq(back[v], after_a[v]) == false then
return ["outcome_differs", [pc[0], tuples[ti], anames[v], after_a[v], back[v], zs_trace(states, parents, vias, ci)]]
end
let v = v + 1
end
let checks = checks + 1
end
else
if fc < 0 then
let after_c = zs_apply_effects(pc[5], c, env_op_c)
let back = zr_retrieved(ec[14], anames, zs_env_add(cconsts, cnames, after_c, ""))
let v = 0
while v < n do
if ze_val_eq(back[v], a[v]) == false then
return ["extra_operation_changes_abstract_state", [pc[0], tuples[ti], anames[v], a[v], back[v], zs_trace(states, parents, vias, ci)]]
end
let v = v + 1
end
let checks = checks + 1
end
end
if ze_rt_error() != "" then
return ["runtime_error", [ze_rt_error()]]
end
let ti = ti + 1
end
let oi = oi + 1
end
let ci = ci + 1
end
let tag = "ok"
if diag[0] == "ok_bounded" then
let tag = "ok_bounded"
end
return [tag, [cn, checks, diag[0] == "ok"]]
end
make a function called zr_derive takes concrete_name, ea returns name
let ec = zs_entry(concrete_name)
let retrieves = ec[14]
let unprimed = []
let primed = []
let cprime = []
let c = 0
let cn = to_num(list_len(ec[1]))
while c < cn do
let cprime = list_push(cprime, [ec[1][c], ["var", ec[1][c] + "'"]])
let c = c + 1
end
let r = 0
let rn = to_num(list_len(retrieves))
while r < rn do
let unprimed = list_push(unprimed, [retrieves[r][0], retrieves[r][2]])
let primed = list_push(primed, [retrieves[r][0] + "'", ze_subst(retrieves[r][2], cprime, [])])
let r = r + 1
end
let both = zs_concat(unprimed, primed)
let facts = []
let f = 0
let fn_total = to_num(list_len(ea[7]))
while f < fn_total do
let facts = list_push(facts, [ea[7][f][0], ea[7][f][1], ze_subst(ea[7][f][2], unprimed, [])])
let f = f + 1
end
let thens = []
let t = 0
let tn = to_num(list_len(ea[12]))
while t < tn do
let thens = list_push(thens, [ea[12][t][0], ea[12][t][1], ze_subst(ea[12][t][2], both, [])])
let t = t + 1
end
let consts = ec[2]
let cnames = zs_pair_names(ec[2])
let k = 0
let kn = to_num(list_len(ea[2]))
while k < kn do
if ze_names_has(cnames, ea[2][k][0]) == false then
let consts = list_appended(consts, ea[2][k])
end
let k = k + 1
end
let derived = concrete_name + "@" + ea[0]
send(zs_registry(), "set", derived, [derived, ec[1], consts, ec[3], ec[4], ec[5], ec[6], facts, ea[8], ea[9], ea[10], ea[11], thens, "", [], ea[15]])
return derived
end
make a function called zs_analyse_feature_refined takes concrete_name, feature_text, max_states returns result
return zs_analyse_feature_refined_ticked(concrete_name, feature_text, max_states, "")
end
make a function called zs_analyse_feature_refined_ticked takes concrete_name, feature_text, max_states, tick_fn returns result
let shape = zr_shape(concrete_name)
if shape[0] != "ok" then
return shape
end
let derived = zr_derive(concrete_name, shape[1])
return zs_analyse_feature_ticked(derived, feature_text, max_states, tick_fn)
end
make a function called zs_format_refinement takes result returns lines
let tag = result[0]
let p = result[1]
if (tag == "ok") or (tag == "ok_bounded") then
let out = ["Refinement holds over " + p[0] + " concrete state(s) (" + p[1] + " instance(s) compared)."]
if tag == "ok_bounded" then
let out = list_push(out, "The search stopped at its bound; states beyond it were not checked.")
end
return out
end
if tag == "init_not_refined" then
return ["The concrete initial state retrieves " + p[0] + " = " + ze_show(p[1]) + ", not the abstract initial value " + ze_show(p[2]) + "."]
end
if tag == "retrieve_breaks_invariant" then
return ["A reachable concrete state retrieves to an abstract state that breaks its invariant: " + p[0]]
end
if tag == "not_applicable" then
return [p[0] + " is enabled on the abstract schema for inputs " + ze_show(p[1]) + ", but the concrete operation's own clause fails: " + p[2]]
end
if tag == "outcome_differs" then
return [p[0] + " with inputs " + ze_show(p[1]) + ": the abstract schema gives " + p[2] + " = " + ze_show(p[3]) + ", the concrete operation retrieves to " + ze_show(p[4]) + "."]
end
if tag == "extra_operation_changes_abstract_state" then
return [p[0] + " has no abstract counterpart, but changes what " + p[2] + " retrieves to: " + ze_show(p[3]) + " before, " + ze_show(p[4]) + " after."]
end
if tag == "operation_missing" then
return ["The concrete schema has no operation named " + p[0] + ", which the abstract schema declares."]
end
if tag == "input_mismatch" then
return [p[0] + " takes " + ze_show(p[2]) + " in the concrete schema and " + ze_show(p[1]) + " in the abstract one."]
end
if tag == "retrieve_incomplete" then
return ["No retrieve line defines the abstract state variable " + p[0] + "."]
end
if tag == "abstract_missing_init" then
return ["The abstract schema has no init for " + p[0] + "."]
end
if tag == "not_a_refinement" then
return [p[0] + " does not declare a refines line."]
end
if tag == "unknown_schema" then
return ["No schema named " + p[0] + " is loaded."]
end
if tag == "unknown_abstract" then
return ["The abstract schema " + p[0] + " named by refines is not loaded."]
end
if tag == "concrete_unsound" then
let out = ["The concrete schema breaks its own invariants, so refinement cannot be checked."]
return zs_concat(out, zs_format_exploration(p[0]))
end
if tag == "runtime_error" then
return ["A clause could not be evaluated: " + p[0]]
end
if tag == "aborted" then
return ["The search was stopped after " + p[0] + " states."]
end
return ["Unrecognised result: " + tag]
end
make a function called zg_bindings_for takes template, input_names, tuple returns pairs
let caps = sm_parse_template(template)[1]
let out = []
let i = 0
let n = to_num(list_len(caps))
while i < n do
let idx = zs_index_of(input_names, caps[i])
let out = list_push(out, [caps[i], tuple[idx]])
let i = i + 1
end
return out
end
make a function called zg_matched_facts takes entry, input_names, tuple, env returns lines
let out = []
let facts = entry[7]
let i = 0
let n = to_num(list_len(facts))
while i < n do
let template = facts[i][0]
if zf_generalised(template, input_names) then
ze_rt_clear()
if ze_eval(facts[i][2], env) == true then
let phrase = sm_fill(template, zg_bindings_for(template, input_names, tuple))
let out = list_push(out, [phrase, sm_parse_template(template)[1]])
end
end
let i = i + 1
end
return out
end
make a function called zg_union_covers takes names, wanted returns ok
let i = 0
let n = to_num(list_len(wanted))
while i < n do
if ze_names_has(names, wanted[i]) == false then
return false
end
let i = i + 1
end
return true
end
make a function called zg_matched_then takes entry, input_names, tuple, env_before, env_after returns phrase
let thens = entry[12]
let i = 0
let n = to_num(list_len(thens))
while i < n do
let template = thens[i][0]
if zf_generalised(template, input_names) then
ze_rt_clear()
if ze_eval(thens[i][2], zs_concat(env_before, env_after)) == true then
return sm_fill(template, zg_bindings_for(template, input_names, tuple))
end
end
let i = i + 1
end
return ""
end
make a function called zg_reject_phrase takes entry, clause_text, input_names, tuple returns phrase
let rejects = entry[15]
let i = 0
let n = to_num(list_len(rejects))
while i < n do
if rejects[i][1] == clause_text then
if zf_generalised(rejects[i][0], input_names) then
return sm_fill(rejects[i][0], zg_bindings_for(rejects[i][0], input_names, tuple))
end
end
let i = i + 1
end
return ""
end
make a function called zg_negative_marker_phrase takes entry, clause_text returns phrase
let negs = entry[9]
if to_num(list_len(negs)) == 0 then
return ""
end
let template = negs[0]
let caps = sm_parse_template(template)[1]
let bindings = []
let i = 0
let n = to_num(list_len(caps))
while i < n do
let bindings = list_push(bindings, [caps[i], clause_text])
let i = i + 1
end
return sm_fill(template, bindings)
end
make a function called zg_find_witness takes states, consts, state_names, prep_op, skip_index returns result
let input_names = prep_op[1]
let requires = prep_op[3]
let tuples = prep_op[2]
let si = 0
let sn = to_num(list_len(states))
while si < sn do
let env_state = zs_env_add(consts, state_names, states[si], "")
let ti = 0
let tn = to_num(list_len(tuples))
while ti < tn do
let env = zs_env_add(env_state, input_names, tuples[ti], "")
let ok = true
let ri = 0
let rn = to_num(list_len(requires))
while ri < rn do
let holds = ze_eval(requires[ri][1], env) == true
if ri == skip_index then
if holds then
let ok = false
end
elif holds == false then
let ok = false
end
let ri = ri + 1
end
if ok then
return ["ok", si, tuples[ti]]
end
let ti = ti + 1
end
let si = si + 1
end
return ["none"]
end
make a function called zg_indent takes lines, keyword_first returns text
let out = sb_new()
let i = 0
let n = to_num(list_len(lines))
while i < n do
if i == 0 then
sb_push(out, " " + keyword_first + " " + lines[i] + "\n")
else
sb_push(out, " And " + lines[i] + "\n")
end
let i = i + 1
end
return sb_str(out)
end
make a function called zg_render_scenario takes entry, op_name, prep_op, states, consts, state_names, state_idx, tuple, violated returns result
let input_names = prep_op[1]
let env = zs_env_add(zs_env_add(consts, state_names, states[state_idx], ""), input_names, tuple, "")
let matched = zg_matched_facts(entry, input_names, tuple, env)
let given = []
let when_lines = []
let covered = []
let m = 0
let mn = to_num(list_len(matched))
while m < mn do
if to_num(list_len(matched[m][1])) == 0 then
if violated == "" then
let given = list_push(given, matched[m][0])
end
else
let when_lines = list_push(when_lines, matched[m][0])
let covered = zs_concat(covered, matched[m][1])
end
let m = m + 1
end
let fully_covered = zg_union_covers(covered, input_names)
if fully_covered == false then
let n = to_num(list_len(input_names))
if n > 0 then
let b = sb_new()
sb_push(b, op_name + " is invoked with ")
let k = 0
while k < n do
if k > 0 then
sb_push(b, ", ")
end
sb_push(b, input_names[k] + " = " + ze_show(tuple[k]))
let k = k + 1
end
let when_lines = list_push(when_lines, sb_str(b))
else
let when_lines = list_push(when_lines, op_name + " is invoked")
end
end
let out = sb_new()
if violated == "" then
sb_push(out, "Scenario: " + op_name + " succeeds\n")
sb_push(out, zg_indent(given, "Given"))
sb_push(out, zg_indent(when_lines, "When"))
let phrase = zg_matched_then(entry, input_names, tuple, env, zs_env_add(entry[2], entry[1], zs_apply_effects(prep_op[5], states[state_idx], env), ""))
if phrase == "" then
let phrase = op_name + " succeeds"
end
sb_push(out, " Then " + phrase + "\n")
else
sb_push(out, "Scenario: " + op_name + " is rejected: " + violated[0] + "\n")
sb_push(out, zg_indent(given, "Given"))
sb_push(out, zg_indent(when_lines, "When"))
let phrase = zg_reject_phrase(entry, violated[0], input_names, tuple)
if phrase == "" then
let phrase = zg_negative_marker_phrase(entry, violated[0])
end
if phrase == "" then
let phrase = op_name + " is rejected because " + violated[0] + " does not hold"
end
sb_push(out, " Then " + phrase + "\n")
end
return [sb_str(out), fully_covered]
end
make a function called zs_generate_feature takes schema_name, op_name, max_states returns result
let entry = zs_entry(schema_name)
if type_of(entry) != "list" then
return ["unknown_schema", [schema_name]]
end
let bfs = zs_bfs(schema_name, max_states, "")
let diag = bfs[0]
if (diag[0] != "ok") and (diag[0] != "ok_bounded") then
return ["schema_unsound", [diag]]
end
let states = bfs[1]
let prep_op = zf_find_prepared(bfs[4], op_name)
if type_of(prep_op) != "list" then
return ["unknown_operation", [schema_name, op_name]]
end
let consts = entry[2]
let state_names = entry[1]
let requires = prep_op[3]
let body = sb_new()
let unwitnessed = []
let normal = zg_find_witness(states, consts, state_names, prep_op, -1)
if normal[0] == "ok" then
sb_push(body, zg_render_scenario(entry, op_name, prep_op, states, consts, state_names, normal[1], normal[2], "")[0])
sb_push(body, "\n")
else
let unwitnessed = list_push(unwitnessed, "(enabled at all)")
end
let ri = 0
let rn = to_num(list_len(requires))
while ri < rn do
let w = zg_find_witness(states, consts, state_names, prep_op, ri)
if w[0] == "ok" then
sb_push(body, zg_render_scenario(entry, op_name, prep_op, states, consts, state_names, w[1], w[2], [requires[ri][0]])[0])
sb_push(body, "\n")
else
let unwitnessed = list_push(unwitnessed, requires[ri][0])
end
let ri = ri + 1
end
let header = sb_new()
sb_push(header, "Feature: " + op_name + " (" + schema_name + ")\n")
let goal_text = zg_goal_text_for(entry, op_name)
if goal_text != "" then
sb_push(header, " Someone wants to " + goal_text + "\n")
end
sb_push(header, "\n")
return ["ok", [sb_str(header) + sb_str(body), unwitnessed]]
end
make a function called zg_goal_text_for takes entry, op_name returns text
let goals = entry[8]
let i = 0
let n = to_num(list_len(goals))
while i < n do
if goals[i][1] == op_name then
return goals[i][0]
end
let i = i + 1
end
return ""
end
make a function called zi_is_digit takes c returns r
return (c >= 48) and (c <= 57)
end
make a function called zi_is_alnum takes c returns r
return zi_is_digit(c) or ((c >= 65) and (c <= 90)) or ((c >= 97) and (c <= 122))
end
make a function called zi_templatize takes text returns result
let h = str_intern(text)
let n = sc_len(h)
let out = sb_new()
let caps = []
let k = 0
let i = 0
while i < n do
let c = sc_code(h, i)
if c == 34 then
let j = i + 1
while (j < n) and (sc_code(h, j) != 34) do
let j = j + 1
end
let inner = sc_substr(h, i + 1, j - i - 1)
let k = k + 1
sb_push(out, "\"{p" + k + "}\"")
let caps = list_push(caps, inner)
let i = j + 1
elif zi_is_alnum(c) then
let j = i
while (j < n) and zi_is_alnum(sc_code(h, j)) do
let j = j + 1
end
let word = sc_substr(h, i, j - i)
if zi_all_digits(word) then
let k = k + 1
sb_push(out, "{p" + k + "}")
let caps = list_push(caps, word)
else
sb_push(out, word)
end
let i = j
else
sb_push(out, sc_char(h, i))
let i = i + 1
end
end
return [sb_str(out), caps]
end
make a function called zi_all_digits takes s returns r
let h = str_intern(s)
let n = sc_len(h)
if n == 0 then
return false
end
let i = 0
while i < n do
if zi_is_digit(sc_code(h, i)) == false then
return false
end
let i = i + 1
end
return true
end
make a function called zi_pascal_case takes text returns name
let h = str_intern(text)
let n = sc_len(h)
let out = sb_new()
let at_start = true
let i = 0
while i < n do
let c = sc_code(h, i)
if zi_is_alnum(c) then
if at_start then
sb_push(out, zi_upper_char(c))
let at_start = false
else
sb_push(out, sc_char(h, i))
end
else
let at_start = true
end
let i = i + 1
end
return sb_str(out)
end
make a function called zi_upper_char takes c returns ch
if (c >= 97) and (c <= 122) then
return chr(c - 32)
end
return chr(c)
end
make a function called zi_join_comma takes items returns text
return zi_join_with(items, ", ")
end
make a function called zi_join_with takes items, sep returns text
let out = sb_new()
let i = 0
let n = to_num(list_len(items))
while i < n do
if i > 0 then
sb_push(out, sep)
end
sb_push(out, items[i])
let i = i + 1
end
return sb_str(out)
end
make a function called zi_add_unique takes list, s returns out
if ze_names_has(list, s) then
return list
end
return list_push(list, s)
end
make a function called zi_infer_skeleton takes feature_text returns text
return zi_infer_skeleton_with_state(feature_text, [])
end
make a function called zi_lower_char takes c returns ch
if (c >= 65) and (c <= 90) then
return chr(c + 32)
end
return chr(c)
end
make a function called zi_lower takes text returns out
let h = str_intern(text)
let n = sc_len(h)
let b = sb_new()
let i = 0
while i < n do
sb_push(b, zi_lower_char(sc_code(h, i)))
let i = i + 1
end
return sb_str(b)
end
make a function called zi_word_to_number takes word returns n
if zi_all_digits(word) then
return to_num(word)
end
let w = zi_lower(word)
if (w == "a") or (w == "an") or (w == "one") then
return 1
end
if w == "two" then
return 2
end
if w == "three" then
return 3
end
if w == "four" then
return 4
end
if w == "five" then
return 5
end
if w == "six" then
return 6
end
if w == "seven" then
return 7
end
if w == "eight" then
return 8
end
if w == "nine" then
return 9
end
if w == "ten" then
return 10
end
return -1
end
make a function called zi_last_word takes text returns word
let h = str_intern(zs_trim(text))
let n = sc_len(h)
let i = n
while (i > 0) and (sc_code(h, i - 1) != 32) do
let i = i - 1
end
return sc_substr(h, i, n - i)
end
make a function called zi_first_word_and_rest takes text returns parts
let h = str_intern(zs_trim(text))
let n = sc_len(h)
let i = 0
while (i < n) and (sc_code(h, i) != 32) do
let i = i + 1
end
return [sc_substr(h, 0, i), zs_trim(sc_substr(h, i, n - i))]
end
make a function called zi_parse_require_phrase takes text returns parsed
let h = str_intern(text)
let n = sc_len(h)
let map_ops = [[" contains at least ", ">="], [" contains at most ", "<="], [" contains exactly ", "=="]]
let k = 0
while k < 3 do
let phrase = map_ops[k][0]
let op = map_ops[k][1]
let idx = sm_find(h, n, phrase, 0)
if idx >= 0 then
let subject = zs_trim(sc_substr(h, 0, idx))
let plen = sc_len(str_intern(phrase))
let rest = zs_trim(sc_substr(h, idx + plen, n - idx - plen))
let qr = zi_first_word_and_rest(rest)
let qty = zi_word_to_number(qr[0])
if (qty >= 0) and (qr[1] != "") then
return ["map", zi_last_word(subject), op, qty, qr[1]]
end
end
let k = k + 1
end
let idx_apos = sm_find(h, n, "'s ", 0)
if idx_apos >= 0 then
let remainder = sc_substr(h, idx_apos + 3, n - idx_apos - 3)
let rh = str_intern(remainder)
let rn = sc_len(rh)
let scalar_ops = [[" is at least ", ">="], [" is at most ", "<="], [" is exactly ", "=="]]
let j = 0
while j < 3 do
let phrase = scalar_ops[j][0]
let op = scalar_ops[j][1]
let idxc = sm_find(rh, rn, phrase, 0)
if idxc >= 0 then
let field = zs_trim(sc_substr(rh, 0, idxc))
let plen = sc_len(str_intern(phrase))
let rest = zs_trim(sc_substr(rh, idxc + plen, rn - idxc - plen))
let qty = zi_word_to_number(zi_first_word_and_rest(rest)[0])
if qty >= 0 then
return ["scalar", field, op, qty]
end
end
let j = j + 1
end
end
return ["none"]
end
make a function called zi_match_require takes text, state_names returns hit
let p = zi_parse_require_phrase(text)
if (p[0] == "map") and ze_names_has(state_names, p[1]) then
return ["hit", "map_get_or(" + p[1] + ", \"" + p[4] + "\", 0) " + p[2] + " " + p[3]]
end
if (p[0] == "scalar") and ze_names_has(state_names, p[1]) then
return ["hit", p[1] + " " + p[2] + " " + p[3]]
end
return ["no"]
end
make a function called zi_parse_ensure_phrase takes text returns parsed
let h = str_intern(text)
let n = sc_len(h)
let map_ops = [[" subtracts ", "-"], [" adds ", "+"]]
let k = 0
while k < 2 do
let phrase = map_ops[k][0]
let sign = map_ops[k][1]
let idx = sm_find(h, n, phrase, 0)
if idx >= 0 then
let subject = zs_trim(sc_substr(h, 0, idx))
let plen = sc_len(str_intern(phrase))
let rest = zs_trim(sc_substr(h, idx + plen, n - idx - plen))
let qr = zi_first_word_and_rest(rest)
let qty = zi_word_to_number(qr[0])
if (qty >= 0) and (qr[1] != "") then
return ["map", zi_last_word(subject), sign, qty, qr[1]]
end
end
let k = k + 1
end
let idx_apos = sm_find(h, n, "'s ", 0)
if idx_apos >= 0 then
let remainder = sc_substr(h, idx_apos + 3, n - idx_apos - 3)
let rh = str_intern(remainder)
let rn = sc_len(rh)
let scalar_ops = [[" is decreased by ", "-"], [" is increased by ", "+"]]
let j = 0
while j < 2 do
let phrase = scalar_ops[j][0]
let sign = scalar_ops[j][1]
let idxc = sm_find(rh, rn, phrase, 0)
if idxc >= 0 then
let field = zs_trim(sc_substr(rh, 0, idxc))
let plen = sc_len(str_intern(phrase))
let rest = zs_trim(sc_substr(rh, idxc + plen, rn - idxc - plen))
let qty = zi_word_to_number(zi_first_word_and_rest(rest)[0])
if qty >= 0 then
return ["scalar", field, sign, qty]
end
end
let j = j + 1
end
let idxb = sm_find(rh, rn, " becomes ", 0)
if idxb >= 0 then
let field = zs_trim(sc_substr(rh, 0, idxb))
let plen = sc_len(str_intern(" becomes "))
let rest = zs_trim(sc_substr(rh, idxb + plen, rn - idxb - plen))
let qty = zi_word_to_number(zi_first_word_and_rest(rest)[0])
if qty >= 0 then
return ["scalar", field, "=", qty]
end
end
end
return ["none"]
end
make a function called zi_match_ensure takes text, state_names returns hit
let p = zi_parse_ensure_phrase(text)
if (p[0] == "map") and ze_names_has(state_names, p[1]) then
return ["hit", p[1] + "' == map_put(" + p[1] + ", \"" + p[4] + "\", map_get(" + p[1] + ", \"" + p[4] + "\") " + p[2] + " " + p[3] + ")"]
end
if (p[0] == "scalar") and ze_names_has(state_names, p[1]) then
if p[2] == "=" then
return ["hit", p[1] + "' == " + p[3]]
end
return ["hit", p[1] + "' == " + p[1] + " " + p[2] + " " + p[3]]
end
return ["no"]
end
make a function called zi_candidate_state_var takes text returns name
let pr = zi_parse_require_phrase(text)
if pr[0] != "none" then
return pr[1]
end
let pe = zi_parse_ensure_phrase(text)
if pe[0] != "none" then
return pe[1]
end
return ""
end
make a function called zi_infer_state_vars takes feature_text returns names
let scenarios = zf_parse_feature(feature_text)[1]
let out = []
let si = 0
let sn = to_num(list_len(scenarios))
while si < sn do
let steps = scenarios[si][1]
let ti = 0
let tn = to_num(list_len(steps))
while ti < tn do
let phase = steps[ti][0]
if (phase == "given") or (phase == "then") then
let cand = zi_candidate_state_var(steps[ti][1])
if cand != "" then
let out = zi_add_unique(out, cand)
end
end
let ti = ti + 1
end
let si = si + 1
end
return out
end
make a function called zi_infer_skeleton_auto takes feature_text returns text
return zi_infer_skeleton_with_state(feature_text, zi_infer_state_vars(feature_text))
end
make a function called zi_infer_skeleton_with_state takes feature_text, state_names returns text
let parsed = zf_parse_feature(feature_text)
let goal_str = parsed[0]
let scenarios = parsed[1]
let op_name = "Op1"
if goal_str != "" then
let op_name = zi_pascal_case(goal_str)
end
let given_templates = []
let when_templates = []
let then_templates = []
let require_clauses = []
let ensure_clauses = []
let si = 0
let sn = to_num(list_len(scenarios))
while si < sn do
let steps = scenarios[si][1]
let when_parts = []
let ti = 0
let tn = to_num(list_len(steps))
while ti < tn do
let phase = steps[ti][0]
if phase == "given" then
let tmpl = zi_templatize(steps[ti][1])[0]
let given_templates = zi_add_unique(given_templates, tmpl)
let req = zi_match_require(steps[ti][1], state_names)
if req[0] == "hit" then
let require_clauses = zi_add_unique(require_clauses, req[1])
end
elif phase == "when" then
let when_parts = list_push(when_parts, steps[ti][1])
else
let tmpl = zi_templatize(steps[ti][1])[0]
let then_templates = zi_add_unique(then_templates, tmpl)
let ens = zi_match_ensure(steps[ti][1], state_names)
if ens[0] == "hit" then
let ensure_clauses = zi_add_unique(ensure_clauses, ens[1])
end
end
let ti = ti + 1
end
if to_num(list_len(when_parts)) > 0 then
let when_tmpl = zi_templatize(zi_join_with(when_parts, " and "))[0]
let when_templates = zi_add_unique(when_templates, when_tmpl)
end
let si = si + 1
end
let op_inputs = []
if to_num(list_len(when_templates)) > 0 then
let op_inputs = sm_parse_template(when_templates[0])[1]
end
let out = sb_new()
let schema_name = "InferredSchema"
if goal_str != "" then
let schema_name = op_name + "Schema"
end
sb_push(out, "schema " + schema_name + "\n")
if to_num(list_len(state_names)) == 0 then
sb_push(out, "# TODO: this schema has no state yet -- Gherkin text names no\n")
sb_push(out, "# state model, so none was guessed at. Add one, e.g.:\n")
sb_push(out, "# state someVariable\n")
sb_push(out, "# init someVariable = ...\n\n")
else
sb_push(out, "state " + zi_join_comma(state_names) + "\n")
sb_push(out, "# TODO: no phrase in the feature says what these start at.\n")
sb_push(out, "# init " + state_names[0] + " = ...\n\n")
end
sb_push(out, "operation " + op_name + "(" + zi_join_comma(op_inputs) + ")\n")
if to_num(list_len(require_clauses)) == 0 then
sb_push(out, " # TODO: replace with the real precondition(s) for " + op_name + ".\n")
sb_push(out, " require true\n")
else
let ri = 0
let rn = to_num(list_len(require_clauses))
while ri < rn do
sb_push(out, " require " + require_clauses[ri] + "\n")
let ri = ri + 1
end
end
if to_num(list_len(ensure_clauses)) == 0 then
sb_push(out, " # TODO: replace with the real effect(s)/postcondition(s).\n")
sb_push(out, " ensure true\n")
else
let ei = 0
let en = to_num(list_len(ensure_clauses))
while ei < en do
sb_push(out, " ensure " + ensure_clauses[ei] + "\n")
let ei = ei + 1
end
end
sb_push(out, "end\n\n")
let gi = 0
let gn = to_num(list_len(given_templates))
while gi < gn do
sb_push(out, "fact " + given_templates[gi] + " => true\n")
let gi = gi + 1
end
if gn > 0 then
sb_push(out, "\n")
end
let wi = 0
let wn = to_num(list_len(when_templates))
while wi < wn do
let caps = sm_parse_template(when_templates[wi])[1]
sb_push(out, "do " + when_templates[wi] + " => " + op_name + "(" + zi_join_comma(caps) + ")\n")
let wi = wi + 1
end
if wn > 0 then
sb_push(out, "\n")
end
let hi = 0
let hn = to_num(list_len(then_templates))
while hi < hn do
sb_push(out, "then " + then_templates[hi] + " => true\n")
let hi = hi + 1
end
if hn > 0 then
sb_push(out, "\n")
end
if goal_str != "" then
sb_push(out, "goal " + goal_str + " => " + op_name + "\n")
end
return sb_str(out)
end
make a function called zi_suggest_requires takes entry, accepted_envs, rejected_envs returns suggestions
let facts = entry[7]
let out = []
let fi = 0
let nfacts = to_num(list_len(facts))
while fi < nfacts do
let fact_text = facts[fi][0]
let fact_ast = facts[fi][2]
let holds_all_accepted = true
let ai = 0
let an = to_num(list_len(accepted_envs))
while ai < an do
ze_rt_clear()
let v = ze_eval(fact_ast, accepted_envs[ai])
if (ze_rt_error() != "") or (v != true) then
let holds_all_accepted = false
end
let ai = ai + 1
end
let fails_some_rejected = false
let counter_examples = 0
let ri = 0
let rn = to_num(list_len(rejected_envs))
while ri < rn do
ze_rt_clear()
let v = ze_eval(fact_ast, rejected_envs[ri])
if (ze_rt_error() != "") or (v != true) then
let fails_some_rejected = true
let counter_examples = counter_examples + 1
end
let ri = ri + 1
end
if holds_all_accepted and (an > 0) and fails_some_rejected then
let out = list_push(out, [fact_text, an, counter_examples, rn])
end
let fi = fi + 1
end
return out
end
make a function called zsdemo_tick takes a, b, c returns go
print("[progress] " + zs_phase() + "|" + a + "|" + b + "|" + c)
return true
end
make a function called zsdemo_lines takes lines returns done
let i = 0
let n = to_num(list_len(lines))
while i < n do
print(lines[i])
let i = i + 1
end
return true
end
make a function called zsdemo_run takes schema_text, feature_text, max_states returns code
zs_set_tick_ms(120)
let loaded = zs_load(schema_text)
if loaded[0] != "ok" then
print("Schema, line " + loaded[1][0] + ": " + loaded[1][1])
print("[done] 2")
return 2
end
let name = loaded[1]
let code = 0
if zs_trim(feature_text) == "" then
let diag = zs_explore_ticked(name, max_states, "zsdemo_tick")
zsdemo_lines(zs_format_exploration(diag))
if (diag[0] != "ok") and (diag[0] != "ok_bounded") then
let code = 1
end
else
let result = zs_analyse_feature_ticked(name, feature_text, max_states, "zsdemo_tick")
zsdemo_lines(zs_format_analysis(result))
if result[0] != "consistent" then
let code = 1
end
end
print("[done] " + code)
return code
end
make a function called zsdemo_generate takes schema_text, op_name, max_states returns code
zs_set_tick_ms(120)
let loaded = zs_load(schema_text)
if loaded[0] != "ok" then
print("Schema, line " + loaded[1][0] + ": " + loaded[1][1])
print("[done] 2")
return 2
end
let name = loaded[1]
zs_set_phase("generating scenarios for " + op_name)
let gen = zs_generate_feature(name, op_name, max_states)
if gen[0] != "ok" then
print("Could not generate scenarios: " + gen[0])
print("[done] 2")
return 2
end
print(gen[1][0])
let unwitnessed = gen[1][1]
let i = 0
let n = to_num(list_len(unwitnessed))
while i < n do
print("(no independent witness for: " + unwitnessed[i] + " -- not given a scenario)")
let i = i + 1
end
print("--- checking the generated feature against the same schema ---")
let checked = zs_analyse_feature_ticked(name, gen[1][0], max_states, "zsdemo_tick")
zsdemo_lines(zs_format_analysis(checked))
let code = 0
if checked[0] != "consistent" then
let code = 1
end
print("[done] " + code)
return code
end
make a function called zsdemo_infer takes feature_text returns code
let skeleton = zi_infer_skeleton_auto(feature_text)
print(skeleton)
print("--- checking that the skeleton is valid, loadable PatLang ---")
let loaded = zs_load(skeleton)
let code = 0
if loaded[0] == "ok" then
let inferred_state = zi_infer_state_vars(feature_text)
if to_num(list_len(inferred_state)) > 0 then
print("Loads cleanly as schema " + loaded[1] + ". State (" + zi_join_comma(inferred_state) + ") and the require/ensure clauses above came from the feature's own wording -- add init values and an input domain before exploring it.")
else
print("Loads cleanly as schema " + loaded[1] + ". Its require/ensure clauses are still the placeholder \"true\" above -- add real state, domains and clauses before exploring it.")
end
else
print("Line " + loaded[1][0] + ": " + loaded[1][1])
let code = 1
end
print("[done] " + code)
return code
end
make a function called zsdemo_refine takes abstract_text, concrete_text, feature_text, max_states returns code
zs_set_tick_ms(120)
let a = zs_load(abstract_text)
if a[0] != "ok" then
print("Abstract schema, line " + a[1][0] + ": " + a[1][1])
print("[done] 2")
return 2
end
let c = zs_load(concrete_text)
if c[0] != "ok" then
print("Concrete schema, line " + c[1][0] + ": " + c[1][1])
print("[done] 2")
return 2
end
let name = c[1]
let result = zs_refines(name, max_states, "zsdemo_tick")
zsdemo_lines(zs_format_refinement(result))
let code = 0
if (result[0] != "ok") and (result[0] != "ok_bounded") then
let code = 1
end
if (code == 0) and (zs_trim(feature_text) != "") then
print("--- carrying a feature already checked against the abstract schema through to the concrete one ---")
let checked = zs_analyse_feature_refined_ticked(name, feature_text, max_states, "zsdemo_tick")
zsdemo_lines(zs_format_analysis(checked))
if checked[0] != "consistent" then
let code = 1
end
end
print("[done] " + code)
return code
end
Timings Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
Times are wall-clock milliseconds or seconds on one Windows 11 machine, taken as the median of three runs unless a row says otherwise. The native column is the project’s pat.exe interpreter, one fresh process per run, and includes reading the libraries. The browser column is headless Chrome running the page’s own WebAssembly interpreter in a background worker, and includes starting the run.
| Live example | Search | Native interpreter | In the browser |
|---|---|---|---|
| LibraryLoans, a return that forgets to lower the count | violation found at step 2 | 0.04 s | 0.71 s |
| LibraryLoans, correct | 44 states, 186 transitions | 0.10 s | 0.77 s |
| Vending machine against listing 15 | 126 states, 388 transitions | 0.16 s | 0.83 s |
| Vending machine that sells with no money | 126 states, 388 transitions | 0.14 s | 0.83 s |
| LibraryLoans, six titles | 668 states, 5,052 transitions | 2.17 s | 3.14 s |
The small examples cost almost the same in the browser whatever they search. A trivial program starts there in about 3 ms, and the 82 KB of libraries with nothing to search takes 0.65 s: reading the libraries is most of every small browser figure. The native interpreter reads them in 0.03 s. Setting that fixed cost aside, the six-title search took 2.49 s in the browser against 2.14 s natively.
How the cost grows was measured on the native interpreter with LibraryLoans at three to seven titles and two members. Time per transition rose by about 33 percent while the state space grew 35-fold:
| Titles | States | Transitions | Time | Per transition |
|---|---|---|---|---|
| 3 | 44 | 186 | 0.06 s | 333 µs |
| 4 | 110 | 600 | 0.22 s | 370 µs |
| 5 | 274 | 1,800 | 0.69 s | 382 µs |
| 6 | 668 | 5,052 | 2.10 s | 415 µs |
| 7 | 1,558 | 13,062 | 5.80 s | 444 µs |
The rise follows the size of each state. A control with a single integer of state, so that a state stays the same size while the number of states doubles, keeps its cost per transition between 34 and 36 µs across a 16-fold increase in states:
| States | Transitions | Time | Per transition |
|---|---|---|---|
| 501 | 500 | 18 ms | 36 µs |
| 1,001 | 1,000 | 34 ms | 34 µs |
| 2,001 | 2,000 | 68 ms | 34 µs |
| 4,001 | 4,000 | 140 ms | 35 µs |
| 8,001 | 8,000 | 282 ms | 35 µs |
At nine titles and two members the search covers 7,060 states and 67,140 transitions. Exploring alone took 34 s, and checking a three-fact feature against the explored states took 51 s in total (single runs, native interpreter). Status queries sent to the command-line driver during the exploration came back in 23 to 68 ms each, including the time to start the second process that sent them.each extra title multiplies the states
Open Questions and Limits Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The checking direction — does this one scenario satisfy the schema — is cheap: substitute concrete values into a predicate and evaluate it, confirmed directly by every example on this page running well under the length of a page load. The harder question — does a feature’s full set of scenarios, taken together across a whole sequence of operations, ever drive the state into something the invariant forbids — is a state-space exploration problem, not a per-scenario evaluation, and it inherits exactly the scaling limit Formal Methods already names for Z on its own: a schema doesn’t check itself past a certain size without tool support. Liu’s process hands that problem to the ProB model checker2. The declared-schema layer takes it on directly, and its cost grows with the state space. The time per transition rose by about a third while the number of states grew thirty-fivefold, and stayed roughly flat when the size of each state was held constant, so the growth comes from each state getting larger and not from the number of states seen so far. A narrower, now-resolved limit: an absence-sentinel ambiguity in pmap_get (a missing key and a genuinely falsy stored value both read back the same way) is a documented hard rule — pmap_has is the only authoritative presence check — rather than a live bug, since every schema built so far uses string- or list-typed state where it doesn’t bite.
References
Liu, B. (2019). Using behavioural specifications to support model-checking [Master’s thesis, University of Waikato]. https://hdl.handle.net/10289/13028 ↩
Liu, B. (2024). Integrating behavioural and formal specifications [PhD thesis, University of Waikato]. https://hdl.handle.net/10289/17330 ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩
Shao, L. (2025). Test generation from Z specifications using LLM support for web system [Unpublished dissertation, University of Waikato]. Supervised by Judy Bowen; shared with the author by Bowen. Not listed in the university’s research repository at the time of writing.
Claessen, K., & Hughes, J. (2000). QuickCheck: A lightweight tool for random testing of Haskell programs. ACM SIGPLAN Notices, 35(9), 268–279. https://doi.org/10.1145/357766.351266 ↩