The Python library

stxt is the STXT parser for Python, published on PyPI: pure Python, no dependencies, 3.10 or later. It implements the five specifications and reports the same error codes as the TypeScript and Java libraries.

The API is that of the other two ports, with snake_case names (getCanonicalName → get_canonical_name). It includes no command line; that of the ecosystem is stxt.

Installation

pip install stxt

The package carries type annotations (py.typed). It only accesses the file system and the environment in the resolution adapters, which can be replaced. Everything public is imported from the root package; the version is in stxt.__version__.

Parsing

Entry point Behaviour
parse_result(text) Collects every error and also returns the nodes it managed to build
parse(text) Raises a ParseException on the first error; with no errors, returns list[Node]
parse_stream(lines) Retains no nodes or errors; see Streaming and parser limits
from stxt import Parser, ParseException

parser = Parser()
result = parser.parse_result(text)

if result.has_errors():
    for error in result.get_errors():
        print(f"line {error.line} [{error.code}]: {error.message}")
roots = result.get_nodes()   # the root nodes, in order; there may be several

# The "raise on first error" form
try:
    nodes = parser.parse(text)
except ParseException as e:
    print(e.line, e.code, e.message)

Every error is a ParseException with line (from 1), code and message, also as get_line(), get_code() and get_message(). Grammar errors are ValidationException, a subclass: a single loop walks both kinds and isinstance tells them apart. After a syntax error the parser recovers and carries on, and the result may include nodes.

The tree

Node is an abstract class with two forms, and InlineNode and TextNode cannot be subclassed:

Class Syntax Own methods
InlineNode Name: value get_value()/set_value(), get_children(), get_child(name), get_children_by_name(name), add_child(), remove_child(), add_inline_node(), add_text_node()
TextNode Name >> get_text_lines(), set_text(), set_text_lines(), add_text_line(), clear_text()

Common methods, in Node:

Method Description
get_name() The name as written
get_canonical_name() The canonical name (STXT-SPEC §4.3)
get_declared_namespace() The namespace written between parentheses, or ""
get_namespace() The effective namespace, inherited from the parents
get_line() The line of the document
get_level() The level, derived from the depth
get_parent() The parent InlineNode, or None at the root
get_text() The value of an inline node, or the joined lines of a block
detach() Detaches the node from its parent

get_children() and get_text_lines() return read-only tuples; the tree is modified with add_child, remove_child and detach.

The form of a node is told apart with isinstance. With the document of the tutorial:

Book (com.acme.book):
	Title: Arquitectura de software moderna
	Authors:
		Author: María Pérez
		Author: Juan García
	ISBN: 978-84-123456-7-8
	Published: 2025-10-01
	Chapter: Introducción
		Content >>
			Conceptos básicos y objetivos del libro.
from stxt import Parser, InlineNode, TextNode, Node

book = Parser().parse_result(text).get_nodes()[0]

book.get_name()            # "Book"
book.get_canonical_name()  # "book"
book.get_namespace()       # "com.acme.book"
book.get_line()            # 1

if isinstance(book, InlineNode):
    book.get_child("Title").get_text()                 # "Arquitectura de software moderna"
    book.get_child("title").get_name()                 # "Title": lookups go by canonical name
    book.get_child("Publisher")                        # None: not there

    authors = book.get_child("Authors")
    [a.get_text() for a in authors.get_children_by_name("Author")]   # ['María Pérez', 'Juan García']
    authors.get_declared_namespace()                   # "": it declares none...
    authors.get_namespace()                            # "com.acme.book": ...it inherits it

    content = book.get_child("Chapter").get_child("Content")
    if isinstance(content, TextNode):
        content.get_text_lines()   # ('Conceptos básicos y objetivos del libro.',)
        content.get_level()        # 2
        content.get_parent() is book.get_child("Chapter")   # True

# A generic walk
def walk(node: Node, depth: int = 0) -> None:
    print("  " * depth + node.get_name())
    if isinstance(node, InlineNode):
        for child in node.get_children():
            walk(child, depth + 1)

Building and editing

Trees are mutable:

  • add_child raises RuntimeException if the node already has a parent (NODE_ALREADY_ATTACHED) or is an ancestor (NODE_CYCLE); remove_child and detach() detach it.
  • set_value and add_text_line refuse a line break (LINE_BREAK_NOT_ALLOWED).
  • The level is derived from the chain of parents, and the line is only set by the parser.
  • In the constructors and factories with two strings, the second one is the content (value or text); the namespace only appears in the three-argument form, InlineNode(name, namespace, value). The keyword arguments value=, namespace= and text= are accepted too.
from stxt import InlineNode, TextNode

email = InlineNode("Email", "com.example.mail", "Weekly report")
email.add_inline_node("From", "Ana García <[email protected]>")
to = email.add_inline_node("To")
to.add_inline_node("Address", "[email protected]")
body = email.add_text_node("Body", "Hi Bob,\n\nSee attached.")

body.get_parent() is email    # True
body.get_level()              # 1
to.get_namespace()            # "com.example.mail", inherited

# Reorder: "To" first
to.detach()
email.add_child(to, 0)

# Edit
email.set_namespace("com.example.docs")   # the whole inheriting subtree follows
body.set_text("Hi Bob,\n\nSee the new attachment.")

Validating against a schema or a template

UnifiedSchemaProvider loads schemas and templates with add_file(text): it parses the definition, validates it against its meta-schema and registers the schema by namespace. SchemaValidator is a Validator that is registered on the Parser and validates each node with a namespace as it is closed.

Template (@stxt.template): com.acme.book
	Structure >>
		Book:
			Title: (1)
			Authors: (1)
				Author: (+)
			ISBN: (1)
			Publisher: (?)
			Published: (?) DATE
			Summary: (?) TEXT
			Chapter: (+)
				Content: (?) TEXT
	Description >>
		Book: Template for publisher book records
from stxt import (
    Parser, UnifiedSchemaProvider, SchemaValidator,
    ValidationException,
)

provider = UnifiedSchemaProvider()
provider.add_file(template_text)            # raises if the template does not validate against its meta-schema

parser = Parser()
parser.register_validator(SchemaValidator(provider))

result = parser.parse_result(document_text)
for error in result.get_errors():
    kind = "schema" if isinstance(error, ValidationException) else "syntax"
    print(f"{kind} line {error.line} [{error.code}]: {error.message}")

With a book without ISBN and with a date that is not YYYY-MM-DD, the loop prints:

schema line 6 [INVALID_VALUE]: Published: Invalid date (1 de octubre de 2025)
schema line 1 [TOO_FEW_CHILDREN]: 0 nodes of 'com.acme.book:isbn' and min is 1
  • A cardinality error is reported on the line of the parent.
  • A namespace the provider does not know produces a SCHEMA_NOT_FOUND per node; the provider does not raise (get_schema() returns None).
  • Nodes without a namespace are not validated (STXT-SCHEMA-SPEC §5). A definition is always validated against its meta-schema.
  • add_file accepts files with several definitions; get_all_schemas() lists them and clear() empties the provider.
  • All the types of STXT-SCHEMA-SPEC §9 are implemented.

The compiled schema can be inspected: provider.get_schema(ns) returns a Schema with get_namespace() and get_node_definition(name); each NodeDefinition has get_type(), get_children() (a dictionary of ChildDefinition by qualified name namespace:name, with get_min() / get_max()), get_values() for an ENUM and get_description().

Grammar resolution

DiscoveryResolver implements STXT-DISCOVERY-SPEC: given the directory of a document, it determines which definitions apply to it (the .stxt/ directories of the document and of its ancestors, then ~/.stxt and /etc/stxt, with per-namespace precedence; STXT_PATH replaces the chain), just like the CLI and the extension.

The resolver receives a DiscoveryFileSystem and a DiscoveryEnvironment. The package includes OsDiscoveryFileSystem (over os) and SystemDiscoveryEnvironment (STXT_PATH, ~/.stxt and /etc/stxt or %ProgramData%\stxt), and the stxt.discovery.resolve(document_dir) shortcut over them; a test can pass an in-memory tree. resolve takes the directory of the document, or None for standard input or an unsaved buffer; in that case the chain starts at the user level. DiscoveryResult implements SchemaProvider and is passed directly to the validator:

from stxt import Parser, SchemaValidator
from stxt.discovery import resolve

discovery = resolve("/home/ana/libros/docs")

discovery.get_chain()        # ['/home/ana/libros/.stxt']  (every ancestor, nearest first)

# Resolution errors are collected, not raised
for error in discovery.get_errors():
    print(f"[{error.code}] {error.file}: {error.message}")

parser = Parser()
parser.register_validator(SchemaValidator(discovery))
result = parser.parse_result(document_text)

DiscoveryResult also records the origin of each definition:

definition = discovery.get_definition("com.acme.book")
definition.file        # '/home/ana/libros/.stxt/@stxt.template/com.acme.book.stxt'
definition.level_dir   # '/home/ana/libros/.stxt'  (the level that won)
definition.schema      # the compiled Schema

discovery.get_active_definitions()   # one per namespace, precedence applied
discovery.get_all_schemas()          # just the schemas of the above

A dedicated DiscoveryResolver caches the levels by directory, to resolve many documents; clear_cache() invalidates them when the definition files may have changed:

from stxt import DiscoveryResolver, OsDiscoveryFileSystem, SystemDiscoveryEnvironment

resolver = DiscoveryResolver(OsDiscoveryFileSystem(), SystemDiscoveryEnvironment())
discovery = resolver.resolve("/home/ana/libros/docs")
resolver.clear_cache()

Errors are DiscoveryError (code, file, message, namespace), with the codes DISCOVERY_DUPLICATE_NAMESPACE, DISCOVERY_NOT_A_DEFINITION, DISCOVERY_NOT_PARSEABLE and DISCOVERY_INVALID_DEFINITION.

The canonical tree as JSON

to_canonical_tree(nodes) returns the JSON value of STXT-TREE-SPEC as a list of dictionaries, the same tree stxt describe emits, and to_canonical_json(nodes) serialises it with two-space indentation. They emit only the normative fields, with no positions or comments.

from stxt import Parser, to_canonical_tree, to_canonical_json

nodes = Parser().parse_result(text).get_nodes()
tree = to_canonical_tree(nodes)
tree[0]["name"]    # 'Book'
tree[0]["form"]    # 'inline'

print(to_canonical_json(nodes))

Writing STXT

NodeWriter serialises a node, or a list of root nodes, in the canonical form of STXT-TREE-SPEC §11, with IndentStyle.TABS (default) or IndentStyle.SPACES_4. It writes the namespace only where it changes from the parent's. Since it starts from the tree, the output has no comments or blank lines outside blocks; it is what stxt format --clean does.

from stxt import NodeWriter, IndentStyle

one = NodeWriter.to_stxt(email)                                          # one node, with tabs
all = NodeWriter.to_stxt_docs(result.get_nodes(), IndentStyle.SPACES_4)   # a whole document

The email built above, written with to_stxt:

Email (com.example.docs): Weekly report
	To:
		Address: [email protected]
	From: Ana García <[email protected]>
	Body >>
		Hi Bob,

		See the new attachment.

Formatter (STXT-TREE-SPEC §12) reformats while keeping comments and blank lines: it rewrites the original text line by line and returns, along with the text, the syntax errors it found. It applies the same rules as stxt format, the extension and the playground.

from stxt import Formatter, IndentStyle

formatted = Formatter.format(source, IndentStyle.TABS)
text, errors = formatted.text, formatted.errors

Observing and validating during the parse

Observer is a base class with four empty callbacks, of which the needed ones are overridden. The parser calls them during the parse: on_create(node, line_string) when a node is opened (already with its parent, effective namespace and level), on_finish(node) when it is closed, on_comment(line_number, line_string) for each comment and on_text_line(node, line_number, line_string, line_indent) for each line of a block.

A Validator runs on each node as it is closed and returns a list of ValidationException, without raising. SchemaValidator is the one the library ships.

from stxt import Parser, Observer, Node

class LoggingObserver(Observer):
    def on_create(self, node: Node, line_string: str) -> None:
        print("open", node.get_qualified_name())
    def on_finish(self, node: Node) -> None:
        print("close", node.get_qualified_name())

parser = Parser()
parser.register_observer(LoggingObserver())
parser.parse_result(text)

Streaming and parser limits

parse_stream(lines) takes an iterable of lines, each without its line break (for instance, an open file), and retains no nodes or errors. The results arrive through a StreamObserver, a base class with two callbacks: on_root_node(node) with each complete, already validated root node, which the parser releases after the call, and on_error(error) with each error, syntax or validation. The memory in use is that of one root tree, which allows processing files that do not fit in memory. A registered StreamObserver receives the same calls with parse and parse_result.

from stxt import Parser, StreamObserver, Node, ParseException

class Counter(StreamObserver):
    def __init__(self) -> None:
        self.roots = 0
    def on_root_node(self, node: Node) -> None:
        self.roots += 1
    def on_error(self, error: ParseException) -> None:
        print(error.line, error.code, error.message)

parser = Parser()
counter = Counter()
parser.register_stream_observer(counter)
with open("log.stxt", encoding="utf-8") as f:
    parser.parse_stream(line.rstrip("\n") for line in f)
counter.roots

The parser applies the limits of STXT-SPEC §11.2, configurable in the constructor: max_nesting (100 levels), max_line_length (10,000 characters) and max_input_size (10,000,000 characters); -1 disables one. Exceeding a limit produces a LimitException (a subclass of ParseException, with the codes LIMIT_NESTING_EXCEEDED, LIMIT_LINE_LENGTH_EXCEEDED and LIMIT_INPUT_SIZE_EXCEEDED) and aborts the parse.

parser = Parser(max_nesting=50, max_input_size=-1)

The API surface

Everything importable from stxt; the subpackages (stxt.schema, stxt.discovery...) are importable too:

Group Names
Parsing Parser, ParseResult, Node, InlineNode, TextNode, NO_LINE, LineIndent, parse_line, EMPTY_NAMESPACE, SPEC_VERSION
Errors ParseException, ValidationException, LimitException, RuntimeException
Extension points Observer, StreamObserver, Validator
Schemas Schema, NodeDefinition, ChildDefinition, SchemaProvider, SchemaProviderMemory, SchemaProviderMeta, SchemaValidator, Type, TypeRegistry, transform_node_to_schema, SCHEMA_NAMESPACE
Templates MetaTemplateSchemaProvider, TemplateSchemaProviderMemory, transform_template_node_to_schema, TEMPLATE_NAMESPACE
Runtime UnifiedSchemaProvider, NodeWriter, IndentStyle, Formatter, FormatResult, to_canonical_tree, to_canonical_json
Resolution DiscoveryResolver, DiscoveryResult, DiscoveryDefinition, DiscoveryLevel, DiscoveryError, DiscoveryFileSystem, DiscoveryEntry, DiscoveryEnvironment, OsDiscoveryFileSystem, SystemDiscoveryEnvironment and stxt.discovery.resolve

The package follows semantic versioning: the API does not change incompatibly within the 1.x line. The status of each specification is in Stability and versions. The changes of each version are in the repository, which is also where errors are reported.