Rye provides two reader-based parsers: Parse-html for HTML and Do-sxml for XML. Both walk the input and call blocks you supply when they encounter matching tags or text. They are useful when you need a few fields from a document rather than a complete in-memory tree.
These examples use the HTML and SXML batteries, so they need a native build with those batteries enabled. They also use reader to turn a string into a reader; for a file, use Reader %page.html or Reader %feed.xml instead. Each recipe is independent.
The selector blocks use <tag> to match an element. A { ... } block supplies nested handlers, while a [ ... ] block runs Rye code on a match. Inside a nested selector, _ [ ... ] handles text. The parser injects the current text or start tag as the value for a handler, so .print prints the matched text.
Suppose a small report has a heading and several paragraphs. Descend through its tags, then handle the text inside the elements you want.
What’s new: reader Parse-html <tag> _
page: "<html><body><h1>Daily report</h1><p>Ready</p><p>Published</p></body></html>"
page |reader |Parse-html {
<html> { <body> {
<h1> { _ [ .print ] }
<p> { _ [ .print ] }
} }
}
This prints Daily report, Ready, and Published on separate lines. For a file, replace page |reader with Reader %report.html. Nesting selectors scopes matches to their parent: the <p> handler above handles paragraphs inside <body>. Text handlers receive text chunks, not necessarily an entire element’s combined text; inline tags may split a sentence into several chunks.
A start-tag handler receives an element, so you can read an attribute without having to parse raw tag strings yourself.
What’s new: Attr? on HTML start tags
page: `<a href="/one">First</a><a href="/two">Second</a>`
page |reader |Parse-html {
<a> [ .Attr? "href" |print ]
}
This prints /one and /two. The returned paths are still relative: the parser does not resolve them against a page URL. Attr? raises an error if a matched tag has no href, so use this recipe when every selected <a> is known to have one. For mixed links, inspect .Attrs? and handle missing keys explicitly.
XML uses Do-sxml, but the nesting and text handlers look similar. This example reads only titles from entries inside a feed.
What’s new: Do-sxml with nested tag handlers
feed: `<feed><entry><title>First</title></entry><entry><title>Second</title></entry></feed>`
feed |reader |Do-sxml {
<feed> { <entry> { <title> { _ [ .print ] } } }
}
It prints First and Second. Replace feed |reader with Reader %feed.xml to read an XML file. A tag handler only sees elements at the selected level: keep the <feed> and <entry> nesting when titles should come specifically from feed entries. XML element names are matched by their local name, without the namespace prefix.
Within a nested XML handler, 'start [ ... ] runs on the opening tag. Use it to read attributes; _ [ ... ] then handles the element’s text.
What’s new: 'start Attr? on XML start elements
feed: `<feed><entry id="one">First</entry><entry id="two">Second</entry></feed>`
feed |reader |Do-sxml {
<feed> { <entry> {
'start [ .Attr? "id" |print ]
_ [ .print ]
} }
}
This prints one, First, two, Second. XML Attr? also accepts an attribute’s local name as a word. Like the HTML variant, it reports an error for a missing attribute; don’t assume optional attributes are present. For namespace-aware handling, inspect .Attrs? before choosing a field.
HTML parsing tolerates everyday HTML markup; XML parsing expects XML structure and exposes XML start-element attributes. Neither example builds a document tree or extracts a fully normalized article. Prefer HTML parsing for pages and XML parsing for feeds or configuration files. For Markdown source, see the Markdown recipes; converting Markdown to HTML is a separate step from parsing HTML.