The Python library
stxt is the STXT parser for Python,
published on PyPI: pure Python, no dependencies, 3.10 or later. It implements the
five specifications and reports the same error codes as the
TypeScript and Java libraries.
The API is that of the other two ports, with snake_case names
(getCanonicalName → get_canonical_name). It includes no command line; that of the
ecosystem is stxt.
Installation
pip install stxt
The package carries type annotations (py.typed). It only accesses the file system
and the environment in the resolution adapters, which can be replaced. Everything
public is imported from the root package; the version is in stxt.__version__.
Parsing
| Entry point | Behaviour |
|---|---|
parse_result(text) |
Collects every error and also returns the nodes it managed to build |
parse(text) |
Raises a ParseException on the first error; with no errors, returns list[Node] |
parse_stream(lines) |
Retains no nodes or errors; see Streaming and parser limits |
from stxt import Parser, ParseException
parser = Parser()
result = parser.parse_result(text)
if result.has_errors():
for error in result.get_errors():
print(f"line {error.line} [{error.code}]: {error.message}")
roots = result.get_nodes() # the root nodes, in order; there may be several
# The "raise on first error" form
try:
nodes = parser.parse(text)
except ParseException as e:
print(e.line, e.code, e.message)
Every error is a ParseException with line (from 1), code and message, also as
get_line(), get_code() and get_message(). Grammar errors are
ValidationException, a subclass: a single loop walks both kinds and isinstance
tells them apart. After a syntax error the parser recovers and carries on, and the
result may include nodes.
The tree
Node is an abstract class with two forms, and InlineNode and TextNode cannot be
subclassed:
| Class | Syntax | Own methods |
|---|---|---|
InlineNode |
Name: value |
get_value()/set_value(), get_children(), get_child(name), get_children_by_name(name), add_child(), remove_child(), add_inline_node(), add_text_node() |
TextNode |
Name >> |
get_text_lines(), set_text(), set_text_lines(), add_text_line(), clear_text() |
Common methods, in Node:
| Method | Description |
|---|---|
get_name() |
The name as written |
get_canonical_name() |
The canonical name (STXT-SPEC §4.3) |
get_declared_namespace() |
The namespace written between parentheses, or "" |
get_namespace() |
The effective namespace, inherited from the parents |
get_line() |
The line of the document |
get_level() |
The level, derived from the depth |
get_parent() |
The parent InlineNode, or None at the root |
get_text() |
The value of an inline node, or the joined lines of a block |
detach() |
Detaches the node from its parent |
get_children() and get_text_lines() return read-only tuples; the tree is modified
with add_child, remove_child and detach.
The form of a node is told apart with isinstance. With the document of the
tutorial:
Book (com.acme.book):
Title: Arquitectura de software moderna
Authors:
Author: María Pérez
Author: Juan García
ISBN: 978-84-123456-7-8
Published: 2025-10-01
Chapter: Introducción
Content >>
Conceptos básicos y objetivos del libro.from stxt import Parser, InlineNode, TextNode, Node
book = Parser().parse_result(text).get_nodes()[0]
book.get_name() # "Book"
book.get_canonical_name() # "book"
book.get_namespace() # "com.acme.book"
book.get_line() # 1
if isinstance(book, InlineNode):
book.get_child("Title").get_text() # "Arquitectura de software moderna"
book.get_child("title").get_name() # "Title": lookups go by canonical name
book.get_child("Publisher") # None: not there
authors = book.get_child("Authors")
[a.get_text() for a in authors.get_children_by_name("Author")] # ['María Pérez', 'Juan García']
authors.get_declared_namespace() # "": it declares none...
authors.get_namespace() # "com.acme.book": ...it inherits it
content = book.get_child("Chapter").get_child("Content")
if isinstance(content, TextNode):
content.get_text_lines() # ('Conceptos básicos y objetivos del libro.',)
content.get_level() # 2
content.get_parent() is book.get_child("Chapter") # True
# A generic walk
def walk(node: Node, depth: int = 0) -> None:
print(" " * depth + node.get_name())
if isinstance(node, InlineNode):
for child in node.get_children():
walk(child, depth + 1)
Building and editing
Trees are mutable:
add_childraisesRuntimeExceptionif the node already has a parent (NODE_ALREADY_ATTACHED) or is an ancestor (NODE_CYCLE);remove_childanddetach()detach it.set_valueandadd_text_linerefuse a line break (LINE_BREAK_NOT_ALLOWED).- The level is derived from the chain of parents, and the line is only set by the parser.
- In the constructors and factories with two strings, the second one is the content
(value or text); the namespace only appears in the three-argument form,
InlineNode(name, namespace, value). The keyword argumentsvalue=,namespace=andtext=are accepted too.
from stxt import InlineNode, TextNode
email = InlineNode("Email", "com.example.mail", "Weekly report")
email.add_inline_node("From", "Ana García <[email protected]>")
to = email.add_inline_node("To")
to.add_inline_node("Address", "[email protected]")
body = email.add_text_node("Body", "Hi Bob,\n\nSee attached.")
body.get_parent() is email # True
body.get_level() # 1
to.get_namespace() # "com.example.mail", inherited
# Reorder: "To" first
to.detach()
email.add_child(to, 0)
# Edit
email.set_namespace("com.example.docs") # the whole inheriting subtree follows
body.set_text("Hi Bob,\n\nSee the new attachment.")
Validating against a schema or a template
UnifiedSchemaProvider loads schemas and templates with add_file(text): it parses
the definition, validates it against its meta-schema and registers the schema by
namespace. SchemaValidator is a Validator that is registered on the Parser and
validates each node with a namespace as it is closed.
Template (@stxt.template): com.acme.book
Structure >>
Book:
Title: (1)
Authors: (1)
Author: (+)
ISBN: (1)
Publisher: (?)
Published: (?) DATE
Summary: (?) TEXT
Chapter: (+)
Content: (?) TEXT
Description >>
Book: Template for publisher book recordsfrom stxt import (
Parser, UnifiedSchemaProvider, SchemaValidator,
ValidationException,
)
provider = UnifiedSchemaProvider()
provider.add_file(template_text) # raises if the template does not validate against its meta-schema
parser = Parser()
parser.register_validator(SchemaValidator(provider))
result = parser.parse_result(document_text)
for error in result.get_errors():
kind = "schema" if isinstance(error, ValidationException) else "syntax"
print(f"{kind} line {error.line} [{error.code}]: {error.message}")
With a book without ISBN and with a date that is not YYYY-MM-DD, the loop prints:
schema line 6 [INVALID_VALUE]: Published: Invalid date (1 de octubre de 2025)
schema line 1 [TOO_FEW_CHILDREN]: 0 nodes of 'com.acme.book:isbn' and min is 1
- A cardinality error is reported on the line of the parent.
- A namespace the provider does not know produces a
SCHEMA_NOT_FOUNDper node; the provider does not raise (get_schema()returnsNone). - Nodes without a namespace are not validated (STXT-SCHEMA-SPEC §5). A definition is always validated against its meta-schema.
add_fileaccepts files with several definitions;get_all_schemas()lists them andclear()empties the provider.- All the types of STXT-SCHEMA-SPEC §9 are implemented.
The compiled schema can be inspected: provider.get_schema(ns) returns a Schema
with get_namespace() and get_node_definition(name); each NodeDefinition has
get_type(), get_children() (a dictionary of ChildDefinition by qualified name
namespace:name, with get_min() / get_max()), get_values() for an ENUM and
get_description().
Grammar resolution
DiscoveryResolver implements STXT-DISCOVERY-SPEC: given the
directory of a document, it determines which definitions apply to it (the .stxt/
directories of the document and of its ancestors, then ~/.stxt and /etc/stxt,
with per-namespace precedence; STXT_PATH replaces the chain), just like the CLI
and the extension.
The resolver receives a DiscoveryFileSystem and a DiscoveryEnvironment. The
package includes OsDiscoveryFileSystem (over os) and
SystemDiscoveryEnvironment (STXT_PATH, ~/.stxt and /etc/stxt or
%ProgramData%\stxt), and the stxt.discovery.resolve(document_dir) shortcut over
them; a test can pass an in-memory tree. resolve takes the directory of the
document, or None for standard input or an unsaved buffer; in that case the chain
starts at the user level. DiscoveryResult implements SchemaProvider and is passed
directly to the validator:
from stxt import Parser, SchemaValidator
from stxt.discovery import resolve
discovery = resolve("/home/ana/libros/docs")
discovery.get_chain() # ['/home/ana/libros/.stxt'] (every ancestor, nearest first)
# Resolution errors are collected, not raised
for error in discovery.get_errors():
print(f"[{error.code}] {error.file}: {error.message}")
parser = Parser()
parser.register_validator(SchemaValidator(discovery))
result = parser.parse_result(document_text)
DiscoveryResult also records the origin of each definition:
definition = discovery.get_definition("com.acme.book")
definition.file # '/home/ana/libros/.stxt/@stxt.template/com.acme.book.stxt'
definition.level_dir # '/home/ana/libros/.stxt' (the level that won)
definition.schema # the compiled Schema
discovery.get_active_definitions() # one per namespace, precedence applied
discovery.get_all_schemas() # just the schemas of the above
A dedicated DiscoveryResolver caches the levels by directory, to resolve many
documents; clear_cache() invalidates them when the definition files may have
changed:
from stxt import DiscoveryResolver, OsDiscoveryFileSystem, SystemDiscoveryEnvironment
resolver = DiscoveryResolver(OsDiscoveryFileSystem(), SystemDiscoveryEnvironment())
discovery = resolver.resolve("/home/ana/libros/docs")
resolver.clear_cache()
Errors are DiscoveryError (code, file, message, namespace), with the codes
DISCOVERY_DUPLICATE_NAMESPACE, DISCOVERY_NOT_A_DEFINITION,
DISCOVERY_NOT_PARSEABLE and DISCOVERY_INVALID_DEFINITION.
The canonical tree as JSON
to_canonical_tree(nodes) returns the JSON value of STXT-TREE-SPEC
as a list of dictionaries, the same tree stxt describe emits, and
to_canonical_json(nodes) serialises it with two-space indentation. They emit only
the normative fields, with no positions or comments.
from stxt import Parser, to_canonical_tree, to_canonical_json
nodes = Parser().parse_result(text).get_nodes()
tree = to_canonical_tree(nodes)
tree[0]["name"] # 'Book'
tree[0]["form"] # 'inline'
print(to_canonical_json(nodes))
Writing STXT
NodeWriter serialises a node, or a list of root nodes, in the canonical form of
STXT-TREE-SPEC §11, with IndentStyle.TABS (default) or
IndentStyle.SPACES_4. It writes the namespace only where it changes from the
parent's. Since it starts from the tree, the output has no comments or blank lines
outside blocks; it is what stxt format --clean does.
from stxt import NodeWriter, IndentStyle
one = NodeWriter.to_stxt(email) # one node, with tabs
all = NodeWriter.to_stxt_docs(result.get_nodes(), IndentStyle.SPACES_4) # a whole document
The email built above, written with to_stxt:
Email (com.example.docs): Weekly report
To:
Address: [email protected]
From: Ana García <[email protected]>
Body >>
Hi Bob,
See the new attachment.Formatter (STXT-TREE-SPEC §12) reformats while keeping
comments and blank lines: it rewrites the original text line by line and returns,
along with the text, the syntax errors it found. It applies the same rules as
stxt format, the extension and the playground.
from stxt import Formatter, IndentStyle
formatted = Formatter.format(source, IndentStyle.TABS)
text, errors = formatted.text, formatted.errors
Observing and validating during the parse
Observer is a base class with four empty callbacks, of which the needed ones are
overridden. The parser calls them during the parse: on_create(node, line_string)
when a node is opened (already with its parent, effective namespace and level),
on_finish(node) when it is closed, on_comment(line_number, line_string) for each
comment and on_text_line(node, line_number, line_string, line_indent) for each
line of a block.
A Validator runs on each node as it is closed and returns a list of
ValidationException, without raising. SchemaValidator is the one the library
ships.
from stxt import Parser, Observer, Node
class LoggingObserver(Observer):
def on_create(self, node: Node, line_string: str) -> None:
print("open", node.get_qualified_name())
def on_finish(self, node: Node) -> None:
print("close", node.get_qualified_name())
parser = Parser()
parser.register_observer(LoggingObserver())
parser.parse_result(text)
Streaming and parser limits
parse_stream(lines) takes an iterable of lines, each without its line break (for
instance, an open file), and retains no nodes or errors. The results arrive through a
StreamObserver, a base class with two callbacks: on_root_node(node) with each
complete, already validated root node, which the parser releases after the call, and
on_error(error) with each error, syntax or validation. The memory in use is that of
one root tree, which allows processing files that do not fit in memory. A registered
StreamObserver receives the same calls with parse and parse_result.
from stxt import Parser, StreamObserver, Node, ParseException
class Counter(StreamObserver):
def __init__(self) -> None:
self.roots = 0
def on_root_node(self, node: Node) -> None:
self.roots += 1
def on_error(self, error: ParseException) -> None:
print(error.line, error.code, error.message)
parser = Parser()
counter = Counter()
parser.register_stream_observer(counter)
with open("log.stxt", encoding="utf-8") as f:
parser.parse_stream(line.rstrip("\n") for line in f)
counter.roots
The parser applies the limits of STXT-SPEC §11.2, configurable
in the constructor: max_nesting (100 levels), max_line_length (10,000 characters)
and max_input_size (10,000,000 characters); -1 disables one. Exceeding a limit
produces a LimitException (a subclass of ParseException, with the codes
LIMIT_NESTING_EXCEEDED, LIMIT_LINE_LENGTH_EXCEEDED and LIMIT_INPUT_SIZE_EXCEEDED)
and aborts the parse.
parser = Parser(max_nesting=50, max_input_size=-1)
The API surface
Everything importable from stxt; the subpackages (stxt.schema,
stxt.discovery...) are importable too:
| Group | Names |
|---|---|
| Parsing | Parser, ParseResult, Node, InlineNode, TextNode, NO_LINE, LineIndent, parse_line, EMPTY_NAMESPACE, SPEC_VERSION |
| Errors | ParseException, ValidationException, LimitException, RuntimeException |
| Extension points | Observer, StreamObserver, Validator |
| Schemas | Schema, NodeDefinition, ChildDefinition, SchemaProvider, SchemaProviderMemory, SchemaProviderMeta, SchemaValidator, Type, TypeRegistry, transform_node_to_schema, SCHEMA_NAMESPACE |
| Templates | MetaTemplateSchemaProvider, TemplateSchemaProviderMemory, transform_template_node_to_schema, TEMPLATE_NAMESPACE |
| Runtime | UnifiedSchemaProvider, NodeWriter, IndentStyle, Formatter, FormatResult, to_canonical_tree, to_canonical_json |
| Resolution | DiscoveryResolver, DiscoveryResult, DiscoveryDefinition, DiscoveryLevel, DiscoveryError, DiscoveryFileSystem, DiscoveryEntry, DiscoveryEnvironment, OsDiscoveryFileSystem, SystemDiscoveryEnvironment and stxt.discovery.resolve |
The package follows semantic versioning: the API does not change incompatibly within the 1.x line. The status of each specification is in Stability and versions. The changes of each version are in the repository, which is also where errors are reported.