The Java library
dev.stxt:stxt-core is the
STXT parser for Java, published on Maven Central. It implements the five
specifications and reports the same error codes as the
TypeScript and Python libraries.
It includes no command line; that of the ecosystem is stxt.
Installation
Requires Java 17 or later and has no runtime dependencies. Under JPMS it is an
automatic module named dev.stxt.
<dependency>
<groupId>dev.stxt</groupId>
<artifactId>stxt-core</artifactId>
<version>1.0.3</version>
</dependency>
Or with Gradle:
implementation 'dev.stxt:stxt-core:1.0.3'
The javadoc is on javadoc.io.
Parsing
| Entry point | Behaviour |
|---|---|
parseResult(text) |
Collects every error and also returns the nodes it managed to build |
parse(text) |
Throws a ParseException on the first error; with no errors, returns List<Node> |
parseStream(reader) |
Retains no nodes or errors; see Streaming and parser limits |
The first two have a file-based counterpart: parseResultFile(File) and
parseFile(File). The STXT facade gives ready-made parsers: STXT.rawParser()
with no validation and STXT.parser(loader) with schema and template validation
registered.
import java.util.List;
import dev.stxt.Node;
import dev.stxt.ParseResult;
import dev.stxt.Parser;
import dev.stxt.exceptions.ParseException;
import dev.stxt.runtime.STXT;
Parser parser = STXT.rawParser();
ParseResult result = parser.parseResult(text);
if (result.hasErrors()) {
for (ParseException error : result.getErrors()) {
System.err.printf("line %d [%s]: %s%n", error.getLine(), error.getCode(), error.getMessage());
}
}
List<Node> roots = result.getNodes(); // the root nodes, in order; there may be several
// The "throw on first error" form
try {
List<Node> nodes = parser.parse(text);
} catch (ParseException e) {
System.err.println(e.getLine() + " " + e.getCode() + " " + e.getMessage());
}
Every error is a ParseException with getLine() (from 1), getCode() and
getMessage(). Grammar errors are ValidationException, a subclass: a single loop
walks both kinds and instanceof tells them apart. After a syntax error the parser
recovers and carries on, and the result may include nodes. The exception hierarchy
is in Errors.
The tree
Node is a sealed class with two forms:
| Class | Syntax | Own methods |
|---|---|---|
InlineNode |
Name: value |
getValue()/setValue(), getChildren(), getChild(name), getChildren(name), addChild(), removeChild(), addInlineNode(), addTextNode() |
TextNode |
Name >> |
getTextLines(), setText(), setTextLines(), addTextLine(), clearText() |
Common methods, in Node:
| Method | Description |
|---|---|
getName() |
The name as written |
getCanonicalName() |
The canonical name (STXT-SPEC §4.3) |
getDeclaredNamespace() |
The namespace written between parentheses, or "" |
getNamespace() |
The effective namespace, inherited from the parents |
getLine() |
The line of the document |
getLevel() |
The level, derived from the depth |
getParent() |
The parent InlineNode, or null at the root |
getText() |
The value of an inline node, or the joined lines of a block |
detach() |
Detaches the node from its parent |
Child lookups go by canonical name. getChild returns null when there is none and
throws AMBIGUOUS_CHILD when there is more than one; repeated children are what
getChildren(name) is for. Both accept a second argument with the namespace.
The form of a node is told apart with instanceof. With the document of the
tutorial:
Book (com.acme.book):
Title: Arquitectura de software moderna
Authors:
Author: María Pérez
Author: Juan García
ISBN: 978-84-123456-7-8
Published: 2025-10-01
Chapter: Introducción
Content >>
Conceptos básicos y objetivos del libro.import dev.stxt.InlineNode;
import dev.stxt.Node;
import dev.stxt.TextNode;
import dev.stxt.runtime.STXT;
Node book = STXT.rawParser().parseResult(text).getNodes().get(0);
book.getName(); // "Book"
book.getCanonicalName(); // "book"
book.getNamespace(); // "com.acme.book"
book.getLine(); // 1
if (book instanceof InlineNode inline) {
inline.getChild("Title").getText(); // "Arquitectura de software moderna"
inline.getChild("title").getName(); // "Title": lookups go by canonical name
inline.getChild("Publisher"); // null: not there
InlineNode authors = (InlineNode) inline.getChild("Authors");
authors.getChildren("Author").stream().map(Node::getText).toList(); // [María Pérez, Juan García]
authors.getDeclaredNamespace(); // "": it declares none...
authors.getNamespace(); // "com.acme.book": ...it inherits it
InlineNode chapter = (InlineNode) inline.getChild("Chapter");
if (chapter.getChild("Content") instanceof TextNode content) {
content.getTextLines(); // [Conceptos básicos y objetivos del libro.]
content.getLevel(); // 2
content.getParent() == chapter; // true
}
}
// A generic walk
static void walk(Node node, int depth) {
System.out.println(" ".repeat(depth) + node.getName());
if (node instanceof InlineNode inline)
for (Node child : inline.getChildren()) walk(child, depth + 1);
}
Building and editing
Trees are mutable:
addChildrefuses a node that already has a parent (NODE_ALREADY_ATTACHED) or that is an ancestor (NODE_CYCLE);removeChildanddetach()detach it. To insert at a position,addChild(index, node).setValueandaddTextLinerefuse a line break (LINE_BREAK_NOT_ALLOWED).- The level is derived from the chain of parents, and the line is only set by the parser.
- In the factories with two
Strings, the second one is the content (value or text); the namespace only appears in the three-argument form.
import dev.stxt.InlineNode;
import dev.stxt.TextNode;
InlineNode email = new InlineNode("Email", "com.example.mail", "Weekly report");
email.addInlineNode("From", "Ana García <[email protected]>");
InlineNode to = email.addInlineNode("To");
to.addInlineNode("Address", "[email protected]");
TextNode body = email.addTextNode("Body", "Hi Bob,\n\nSee attached.");
body.getParent() == email; // true
body.getLevel(); // 1
to.getNamespace(); // "com.example.mail", inherited
// Reorder: "To" first
to.detach();
email.addChild(0, to);
// Edit
email.setNamespace("com.example.docs"); // the whole inheriting subtree follows
body.setText("Hi Bob,\n\nSee the new attachment.");
Validating against a schema or a template
A ResourcesLoader determines where schemas and templates come from.
ResourcesLoaderDirectory looks them up on disk as
<dir>/@stxt.schema/<namespace>.stxt and <dir>/@stxt.template/<namespace>.stxt,
the layout stxt install leaves. STXT.parser(loader) returns a parser that
validates each definition against its meta-schema when loading it, caches it and
validates every node with a namespace as it is closed.
Template (@stxt.template): com.acme.book
Structure >>
Book:
Title: (1)
Authors: (1)
Author: (+)
ISBN: (1)
Publisher: (?)
Published: (?) DATE
Summary: (?) TEXT
Chapter: (+)
Content: (?) TEXT
Description >>
Book: Template for publisher book recordsimport java.io.File;
import dev.stxt.ParseResult;
import dev.stxt.Parser;
import dev.stxt.exceptions.ParseException;
import dev.stxt.exceptions.ValidationException;
import dev.stxt.resources.ResourcesLoader;
import dev.stxt.resources.ResourcesLoaderDirectory;
import dev.stxt.runtime.STXT;
ResourcesLoader loader = new ResourcesLoaderDirectory(new File("/home/ana/libros/.stxt"));
Parser parser = STXT.parser(loader);
ParseResult result = parser.parseResult(documentText);
for (ParseException error : result.getErrors()) {
String kind = (error instanceof ValidationException) ? "schema" : "syntax";
System.out.printf("%s line %d [%s]: %s%n", kind, error.getLine(), error.getCode(), error.getMessage());
}
With a book without ISBN and with a date that is not YYYY-MM-DD, the loop prints:
schema line 6 [INVALID_VALUE]: Published: Invalid date (1 de octubre de 2025)
schema line 1 [TOO_FEW_CHILDREN]: 0 nodes of 'com.acme.book:isbn' and min is 1
ResourcesLoader is a single-method interface, retrieve(namespace, resource),
where namespace is @stxt.schema or @stxt.template and resource the
namespace being looked up. A definition in memory, on the classpath or in a database
is served with a lambda that returns its text, or throws ResourceNotFoundException:
import dev.stxt.exceptions.ResourceNotFoundException;
ResourcesLoader loader = (namespace, resource) -> {
if (namespace.equals("@stxt.template") && resource.equals("com.acme.book")) return templateText;
throw new ResourceNotFoundException(namespace, resource);
};
Parser parser = STXT.parser(loader);
- A cardinality error is reported on the line of the parent.
- A namespace the loader does not know produces a
SCHEMA_NOT_FOUNDper node; the provider does not throw (getSchema()returnsnull). - A definition that does not validate against its meta-schema is reported as a
finding of the document (for instance
TYPE_NOT_VALID) and not as an exception. - Nodes without a namespace are not validated (STXT-SCHEMA-SPEC §5). A definition is always validated against its meta-schema.
- All the types of STXT-SCHEMA-SPEC §9 are implemented.
The facade is equivalent to registering a SchemaValidator over its provider:
import dev.stxt.schema.SchemaValidator;
Parser parser = new Parser();
parser.registerValidator(new SchemaValidator(STXT.schemaProvider(loader)));
The compiled schema can be inspected: STXT.schemaProvider(loader) is the
SchemaProvider the facade uses, and getSchema(ns) returns a Schema with
getNamespace(), getNodes() and getNodeDefinition(name); each
NodeDefinition has getType(), getChildren() (a map of ChildDefinition by
qualified name namespace:name, with getMin() / getMax()), getValues() for
an ENUM and getDescription().
Grammar resolution
ResourcesLoaderDirectory looks in a single directory. DiscoveryResolver
implements STXT-DISCOVERY-SPEC: given the directory of a
document, it determines which definitions apply to it (the .stxt/ directories of
the document and of its ancestors, then ~/.stxt and /etc/stxt, with
per-namespace precedence; STXT_PATH replaces the chain), just like the CLI and the
extension.
The no-argument constructor uses NioDiscoveryFileSystem (over java.nio.file) and
SystemDiscoveryEnvironment (STXT_PATH, user and system directories); both can be
replaced in a test (DiscoveryFileSystem, DiscoveryEnvironment). resolve takes
the directory of the document, or null for standard input or an unsaved buffer; in
that case the chain starts at the user level. DiscoveryResult implements
SchemaProvider and is passed directly to the validator:
import java.nio.file.Path;
import dev.stxt.Parser;
import dev.stxt.discovery.DiscoveryDefinition;
import dev.stxt.discovery.DiscoveryError;
import dev.stxt.discovery.DiscoveryResolver;
import dev.stxt.discovery.DiscoveryResult;
import dev.stxt.schema.SchemaValidator;
DiscoveryResolver resolver = new DiscoveryResolver();
DiscoveryResult discovery = resolver.resolve(Path.of("/home/ana/libros/docs"));
discovery.getChain(); // [/home/ana/libros/.stxt] (every ancestor, nearest first)
// Resolution errors are collected, not thrown
for (DiscoveryError error : discovery.getErrors()) {
System.err.printf("[%s] %s: %s%n", error.getCode(), error.getFile(), error.getMessage());
}
Parser parser = new Parser();
parser.registerValidator(new SchemaValidator(discovery));
ParseResult result = parser.parseResult(documentText);
DiscoveryResult also records the origin of each definition:
DiscoveryDefinition definition = discovery.getDefinition("com.acme.book");
definition.getFile(); // /home/ana/libros/.stxt/@stxt.template/com.acme.book.stxt
definition.getLevelDir(); // /home/ana/libros/.stxt (the level that won)
definition.getSchema(); // the compiled Schema
discovery.getActiveDefinitions(); // one per namespace, precedence applied
discovery.getAllSchemas(); // just the schemas of the above
Levels are cached by directory; resolver.clearCache() invalidates them when the
definition files may have changed. Errors are DiscoveryError (getCode(),
getFile(), getMessage(), getNamespace()), with the codes
DISCOVERY_DUPLICATE_NAMESPACE, DISCOVERY_NOT_A_DEFINITION,
DISCOVERY_NOT_PARSEABLE and DISCOVERY_INVALID_DEFINITION.
The canonical tree as JSON
TreeJson.toCanonicalJson(nodes), or toCanonicalJson(node) for a single root,
returns as text the JSON of STXT-TREE-SPEC, the same one
stxt describe emits, with two-space indentation. It emits only the normative
fields, with no positions or comments, and depends on no JSON library.
import dev.stxt.runtime.TreeJson;
String json = TreeJson.toCanonicalJson(result.getNodes());
System.out.print(json);
Writing STXT
NodeWriter serialises a node, or a list of root nodes, in the canonical form of
STXT-TREE-SPEC §11, with IndentStyle.TABS (default) or
IndentStyle.SPACES_4. It writes the namespace only where it changes from the
parent's. Since it starts from the tree, the output has no comments or blank lines
outside blocks; it is what stxt format --clean does.
import dev.stxt.runtime.NodeWriter;
import dev.stxt.runtime.NodeWriter.IndentStyle;
String one = NodeWriter.toSTXT(email); // one node, with tabs
String all = NodeWriter.toSTXT(result.getNodes(), IndentStyle.SPACES_4); // a whole document
The email built above, written with toSTXT:
Email (com.example.docs): Weekly report
To:
Address: [email protected]
From: Ana García <[email protected]>
Body >>
Hi Bob,
See the new attachment.Formatter (STXT-TREE-SPEC §12) reformats while keeping
comments and blank lines: it rewrites the original text line by line and returns,
along with the text, the syntax errors it found. It applies the same rules as
stxt format, the extension and the playground.
import java.util.List;
import dev.stxt.exceptions.ParseException;
import dev.stxt.runtime.Formatter;
import dev.stxt.runtime.FormatResult;
import dev.stxt.runtime.NodeWriter.IndentStyle;
FormatResult formatted = Formatter.format(source, IndentStyle.TABS);
String text = formatted.text();
List<ParseException> errors = formatted.errors();
Observing and validating during the parse
A registered Observer receives calls during the parse: onCreate(node, line) when
a node is opened (already with its parent, effective namespace and level),
onFinish(node) when it is closed, onComment(lineNumber, line) for each comment
and onTextLine(node, lineNumber, lineString, line) for each line of a block.
A Validator runs on each node as it is closed and returns a list of
ValidationException, without throwing. SchemaValidator is the one the library
ships; Validator is a functional interface and accepts a lambda.
import java.util.List;
import dev.stxt.LineIndent;
import dev.stxt.Node;
import dev.stxt.Parser;
import dev.stxt.TextNode;
import dev.stxt.exceptions.ValidationException;
import dev.stxt.processors.Observer;
Parser parser = new Parser();
parser.registerObserver(new Observer() {
@Override public void onCreate(Node node, String line) { System.out.println("open " + node.getQualifiedName()); }
@Override public void onFinish(Node node) { System.out.println("close " + node.getQualifiedName()); }
@Override public void onComment(int lineNumber, String line) { }
@Override public void onTextLine(TextNode node, int lineNumber, String lineString, LineIndent line) { }
});
parser.registerValidator(node -> List.<ValidationException>of()); // a validator that never complains
parser.parseResult(text);
Streaming and parser limits
parseStream(reader) reads from a Reader and retains no nodes or errors. The
results arrive through a StreamObserver: onRootNode(node) with each complete,
already validated root node, which the parser releases after the call, and
onError(error) with each error, syntax or validation. The memory in use is that of
one root tree, which allows processing files that do not fit in memory. A registered
StreamObserver receives the same calls with parse and parseResult.
import java.io.Reader;
import java.nio.file.Files;
import java.nio.file.Path;
import dev.stxt.Node;
import dev.stxt.Parser;
import dev.stxt.exceptions.ParseException;
import dev.stxt.processors.StreamObserver;
Parser parser = new Parser();
parser.registerStreamObserver(new StreamObserver() {
@Override public void onRootNode(Node node) { System.out.println(node.getQualifiedName()); }
@Override public void onError(ParseException error) { System.err.println(error.getLine() + " " + error.getCode()); }
});
try (Reader reader = Files.newBufferedReader(Path.of("log.stxt"))) {
parser.parseStream(reader);
}
The parser applies the limits of STXT-SPEC §11.2, configurable
with setMaxNesting (100 levels), setMaxLineLength (10,000 characters) and
setMaxInputSize (10,000,000 characters); -1 disables one. Exceeding a limit
produces a LimitException (a subclass of ParseException, with the codes
LIMIT_NESTING_EXCEEDED, LIMIT_LINE_LENGTH_EXCEEDED and LIMIT_INPUT_SIZE_EXCEEDED)
and aborts the parse.
Parser parser = new Parser();
parser.setMaxNesting(50);
parser.setMaxInputSize(-1);
Errors
Every exception is unchecked and extends dev.stxt.exceptions.STXTException, with
its code in getCode():
| Exception | When |
|---|---|
ParseException |
The syntax is wrong; adds getLine() |
ValidationException |
The document breaks its grammar (type, cardinality, undeclared child); a subclass of ParseException |
LimitException |
A parser limit was exceeded (LIMIT_*); a subclass of ParseException |
SchemaException |
The schema or template itself is malformed |
ResourceNotFoundException |
A ResourcesLoader has no such resource (providers turn it into a SCHEMA_NOT_FOUND finding) |
STXTIOException |
Reading a file failed |
STXTException (base) |
Tree integrity (NODE_ALREADY_ATTACHED, NODE_CYCLE, LINE_BREAK_NOT_ALLOWED), an ambiguous lookup (AMBIGUOUS_CHILD) and other runtime failures |
Those of parseResult are not thrown: they are collected in getErrors().
Tree-integrity ones are thrown.
The API surface
| Package | What is there |
|---|---|
dev.stxt |
Parser, ParseResult, Node, InlineNode, TextNode, Constants, LineIndent, LineIndentParser, NameNamespace, NameNamespaceParser, NamespaceValidator |
dev.stxt.exceptions |
STXTException, ParseException, ValidationException, LimitException, SchemaException, ResourceNotFoundException, STXTIOException |
dev.stxt.processors |
Observer, StreamObserver, Validator |
dev.stxt.schema |
Schema, NodeDefinition, ChildDefinition, SchemaProvider, SchemaValidator, SchemaProviderResources, SchemaProviderCache, SchemaProviderMemory, SchemaProviderMeta, SchemaParser, DefinitionCompiler, Type, TypeRegistry, and the types in dev.stxt.schema.type |
dev.stxt.template |
TemplateParser, TemplateSchemaProvider, TemplateSchemaProviderMemory, MetaTemplateSchemaProvider |
dev.stxt.resources |
ResourcesLoader, ResourcesLoaderDirectory |
dev.stxt.runtime |
STXT (the facade), UnifiedSchemaProvider, NodeWriter and IndentStyle, Formatter and FormatResult, TreeJson |
dev.stxt.discovery |
DiscoveryResolver, DiscoveryResult, DiscoveryDefinition, DiscoveryLevel, DiscoveryError, DiscoveryEnvironment, SystemDiscoveryEnvironment, DiscoveryFileSystem, NioDiscoveryFileSystem |
The package follows semantic versioning: the API does not change incompatibly within
the 1.x line. The status of each specification is in
Stability and versions. The changes of each version are in the
CHANGELOG.md of the repository, which is
also where errors are reported.