orgparse — Python module for reading Emacs org-mode files¶
Install¶
You can install orgparse via PyPI:
pip install orgparse
Or via conda-forge:
conda install orgparse -c conda-forge
Usage¶
The API documentation includes extensive doctests for individual methods. Here are some examples to get started.
Load an org document¶
from orgparse import load, loads
load('PATH/TO/FILE.org')
load(file_like_object)
loads('''
* This is org-mode contents
You can load org object from string.
** Second header
''')
See the loading implementation.
Traverse an org tree¶
>>> from orgparse import loads
>>> root = loads('''
... * Heading 1
... ** Heading 2
... *** Heading 3
... ''')
>>> for node in root[1:]: # [1:] for skipping root itself
... print(node)
* Heading 1
** Heading 2
*** Heading 3
>>> h1 = root.children[0]
>>> h2 = h1.children[0]
>>> h3 = h2.children[0]
>>> print(h1)
* Heading 1
>>> print(h2)
** Heading 2
>>> print(h3)
*** Heading 3
>>> print(h2.get_parent())
* Heading 1
>>> print(h3.get_parent(max_level=1))
* Heading 1
Access node attributes¶
>>> root = loads('''
... * DONE Heading :TAG:
... CLOSED: [2012-02-26 Sun 21:15] SCHEDULED: <2012-02-26 Sun>
... CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05
... :PROPERTIES:
... :Effort: 1:00
... :OtherProperty: some text
... :END:
... Body texts...
... ''')
>>> node = root.children[0]
>>> node.heading
'Heading'
>>> node.scheduled
OrgDateScheduled((2012, 2, 26))
>>> node.closed
OrgDateClosed((2012, 2, 26, 21, 15, 0))
>>> node.clock
[OrgDateClock((2012, 2, 26, 21, 10, 0), (2012, 2, 26, 21, 15, 0))]
>>> bool(node.deadline) # it is not specified
False
>>> node.tags == set(['TAG'])
True
>>> node.get_property('Effort')
60
>>> node.get_property('UndefinedProperty') # returns None
>>> node.get_property('OtherProperty')
'some text'
>>> node.body
' Body texts...'
Read named tables¶
Tables in node.body_rich expose their #+NAME: through Table.name.
The name is None for unnamed tables.
>>> from orgparse.extra import Table
>>> root = loads('''
... #+NAME: measurements
... | x | y |
... |---+---|
... | 1 | 2 |
... ''')
>>> [table] = [part for part in root.body_rich if isinstance(part, Table) and part.name == 'measurements']
>>> list(table.as_dicts)
[{'x': '1', 'y': '2'}]
More examples¶
The tests show additional supported features:
Development and documentation¶
Clone with git clone --recurse-submodules https://github.com/karlicoss/orgparse.git to include the test corpus.
For an existing checkout, run git submodule update --init --recursive.
Run the tests with uv tool run --with tox-uv tox -e tests.
This also checks the examples in this README.
Corpus tests parse upstream Org examples, Sacha Chua’s Emacs configuration, exobrain notes, and nvim-orgmode documents from testdata/external.
They check tree structure, source line ranges, and attribute access.
Edit README.qmd, then regenerate README.md with uv tool run --with tox-uv tox -e quarto.
Quarto computes links to source code and tests from their definitions, so line numbers are refreshed when rendering.
Commit both files together; CI checks that the generated README is current.
Build the documentation with uv tool run --with tox-uv tox -e docs and open doc/_build/html/index.html.
Sphinx combines the generated README with the API docstrings; it does not need Quarto to build the site.
Read the Docs uses .readthedocs.yaml to run the same docs environment.
The latest version follows master, and stable follows releases.
Automatic builds require the GitHub integration in the Read the Docs project settings.
Project status¶
The project is maintained by @karlicoss.
For my personal use, orgparse mostly has all features I need, so there hasn’t been much active development lately.
However, contributions are always welcome! Please provide tests along with your contribution if you’re fixing bugs or adding new functionality.
Loading documents¶
Read Emacs org-mode files as trees of Python objects.
- class orgparse.OrgBaseNode(env: OrgEnv, index: int | None = None)[source]¶
Base class for
OrgRootNodeandOrgNode- env
An instance of
OrgEnv. All nodes in a same file shares same instance.
OrgBaseNodeis an iterable object:>>> from orgparse import loads >>> root = loads(''' ... * Heading 1 ... ** Heading 2 ... *** Heading 3 ... * Heading 4 ... ''') >>> for node in root: ... print(node) * Heading 1 ** Heading 2 *** Heading 3 * Heading 4
Note that the first blank line is due to the root node, as iteration contains the object itself. To skip that, use slice access
[1:]:>>> for node in root[1:]: ... print(node) * Heading 1 ** Heading 2 *** Heading 3 * Heading 4
It also supports sequence protocol.
>>> print(root[1]) * Heading 1 >>> root[0] is root # index 0 means itself True >>> len(root) # remember, sequence contains itself 5
Note the difference between
root[1:]androot[1]:>>> for node in root[1]: ... print(node) * Heading 1 ** Heading 2 *** Heading 3
Nodes remember the line number information (1-indexed):
>>> print(root.children[1].linenumber) 5
- property end_linenumber: int[source]¶
One-based, inclusive end line of this node’s own source text.
Includes the heading, metadata, and trailing blank lines, excluding descendants. For
OrgRootNode, this covers the preamble before the first heading and returns 0 when it is empty. Usenode[-1].end_linenumberfor the end of the entire subtree.>>> from orgparse import loads >>> root = loads('''Preamble ... * Parent ... body ... ** Child ... child body ... * Last''') >>> [(node.linenumber, node.end_linenumber) for node in root] [(1, 1), (2, 3), (4, 5), (6, 6)] >>> root[1][-1].end_linenumber 5 >>> loads('* Heading').end_linenumber 0
- property previous_same_level: OrgBaseNode | None[source]¶
Return previous node if exists or None otherwise.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... * Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.previous_same_level is None True >>> n2.previous_same_level is n1 True >>> n3.previous_same_level is None # n2 is not at the same level True
- property next_same_level: OrgBaseNode | None[source]¶
Return next node if exists or None otherwise.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... * Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.next_same_level is n2 True >>> n2.next_same_level is None # n3 is not at the same level True >>> n3.next_same_level is None True
- get_parent(max_level: int | None = None)[source]¶
Return a parent node.
- Parameters:
max_level (int) –
In the normally structured org file, it is a level of the ancestor node to return. For example,
get_parent(max_level=0)returns a root node.In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to
max_level, this function try to find an ancestor node which level is smaller thanmax_level.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
For simplicity, accessing
parentis alias of callingget_parent()without argument.>>> n1.get_parent() is n1.parent True >>> root.parent is None True
This is a little bit pathological situation – but works.
>>> root = loads(''' ... * Node 1 ... *** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
Now let’s play with max_level.
>>> root = loads(''' ... * Node 1 (level 1) ... ** Node 2 (level 2) ... *** Node 3 (level 3) ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n3.get_parent() is n2 True >>> n3.get_parent(max_level=2) is n2 # same as default True >>> n3.get_parent(max_level=1) is n1 True >>> n3.get_parent(max_level=0) is root True
- property parent[source]¶
Alias of
get_parent()(calling without argument).
- property children[source]¶
A list of child nodes.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... *** Node 3 ... ** Node 4 ... ''') >>> (n1, n2, n3, n4) = list(root[1:]) >>> (c1, c2) = n1.children >>> c1 is n2 True >>> c2 is n4 True
Note the difference to
n1[1:], which returns the Node 3 also:>>> (m1, m2, m3) = list(n1[1:]) >>> m2 is n3 True
- property root[source]¶
The root node.
>>> from orgparse import loads >>> root = loads('* Node 1') >>> n1 = root[1] >>> n1.root is root True
- property properties: dict[str, str | int | float][source]¶
Node properties as a dictionary.
>>> from orgparse import loads >>> root = loads(''' ... * Node ... :PROPERTIES: ... :SomeProperty: value ... :END: ... ''') >>> root.children[0].properties['SomeProperty'] 'value'
- get_property(key, val=None) str | int | float | None[source]¶
Return property named
keyif exists orvalotherwise.- Parameters:
key (str) – Key of property.
val – Default value to return.
- property tags: set[str][source]¶
Tags of this and parent’s node.
>>> from orgparse import loads >>> n2 = loads(''' ... * Node 1 :TAG1: ... ** Node 2 :TAG2: ... ''')[2] >>> n2.tags == set(['TAG1', 'TAG2']) True
- property shallow_tags: set[str][source]¶
Tags defined for this node (don’t look-up parent nodes).
>>> from orgparse import loads >>> n2 = loads(''' ... * Node 1 :TAG1: ... ** Node 2 :TAG2: ... ''')[2] >>> n2.shallow_tags == set(['TAG2']) True
- is_root()[source]¶
Return
Truewhen it is a root node.>>> from orgparse import loads >>> root = loads('* Node 1') >>> root.is_root() True >>> n1 = root[1] >>> n1.is_root() False
- get_timestamps(active=False, inactive=False, range=False, point=False)[source]¶
Return a list of timestamps in the body text.
- Parameters:
- Return type:
list of
orgparse.date.OrgDatesubclasses
Consider the following org node:
>>> from orgparse import loads >>> node = loads(''' ... * Node ... CLOSED: [2012-02-26 Sun 21:15] SCHEDULED: <2012-02-26 Sun> ... CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05 ... Some inactive timestamp [2012-02-23 Thu] in body text. ... Some active timestamp <2012-02-24 Fri> in body text. ... Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon]. ... Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>. ... ''').children[0]
The default flags are all off, so it does not return anything.
>>> node.get_timestamps() []
You can fetch appropriate timestamps using keyword arguments.
>>> node.get_timestamps(inactive=True, point=True) [OrgDate((2012, 2, 23), None, False)] >>> node.get_timestamps(active=True, point=True) [OrgDate((2012, 2, 24))] >>> node.get_timestamps(inactive=True, range=True) [OrgDate((2012, 2, 25), (2012, 2, 27), False)] >>> node.get_timestamps(active=True, range=True) [OrgDate((2012, 2, 26), (2012, 2, 28))]
This is more complex example. Only active timestamps, regardless of range/point type.
>>> node.get_timestamps(active=True, point=True, range=True) [OrgDate((2012, 2, 24)), OrgDate((2012, 2, 26), (2012, 2, 28))]
- property datelist[source]¶
Alias of
.get_timestamps(active=True, inactive=True, point=True).- Return type:
list of
orgparse.date.OrgDatesubclasses
>>> from orgparse import loads >>> root = loads(''' ... * Node with point dates <2012-02-25 Sat> ... CLOSED: [2012-02-25 Sat 21:15] ... Some inactive timestamp [2012-02-26 Sun] in body text. ... Some active timestamp <2012-02-27 Mon> in body text. ... ''') >>> root.children[0].datelist [OrgDate((2012, 2, 25)), OrgDate((2012, 2, 26), None, False), OrgDate((2012, 2, 27))]
- property rangelist[source]¶
Alias of
.get_timestamps(active=True, inactive=True, range=True).- Return type:
list of
orgparse.date.OrgDatesubclasses
>>> from orgparse import loads >>> root = loads(''' ... * Node with range dates <2012-02-25 Sat>--<2012-02-28 Tue> ... CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05 ... Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon]. ... Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>. ... Some time interval <2012-02-27 Mon 11:23-12:10>. ... ''') >>> root.children[0].rangelist [OrgDate((2012, 2, 25), (2012, 2, 28)), OrgDate((2012, 2, 25), (2012, 2, 27), False), OrgDate((2012, 2, 26), (2012, 2, 28)), OrgDate((2012, 2, 27, 11, 23, 0), (2012, 2, 27, 12, 10, 0))]
- class orgparse.OrgEnv(todos: Sequence[str] | None = None, dones: Sequence[str] | None = None, filename: str | Path = '<undefined>')[source]¶
Information global to the file (e.g, TODO keywords).
- property nodes: list[OrgBaseNode][source]¶
A list of org nodes.
>>> OrgEnv().nodes # default is empty (of course) []
>>> from orgparse import loads >>> loads(''' ... * Heading 1 ... ** Heading 2 ... *** Heading 3 ... ''').env.nodes [<orgparse.node.OrgRootNode object at 0x...>, <orgparse.node.OrgNode object at 0x...>, <orgparse.node.OrgNode object at 0x...>, <orgparse.node.OrgNode object at 0x...>]
- property todo_keys[source]¶
TODO keywords defined for this document (file).
>>> env = OrgEnv() >>> env.todo_keys ['TODO']
- property done_keys[source]¶
DONE keywords defined for this document (file).
>>> env = OrgEnv() >>> env.done_keys ['DONE']
- property all_todo_keys[source]¶
All TODO keywords (including DONEs).
>>> env = OrgEnv() >>> env.all_todo_keys ['TODO', 'DONE']
- property filename: str[source]¶
Return the source filename as a string.
A
pathlib.Pathpassed toOrgEnvis converted to a string. Documents loaded without a filename use a placeholder such as<string>.
- class orgparse.OrgNode(*args, **kwds)[source]¶
Node to represent normal org node
See
OrgBaseNodefor other available functions.- get_heading(format: str = 'plain') str[source]¶
Return a string of head text without tags and TODO keywords.
>>> from orgparse import loads >>> node = loads('* TODO Node 1').children[0] >>> node.get_heading() 'Node 1'
It strips off inline markup by default (
format='plain'). You can get the original raw string by specifyingformat='raw'.>>> node = loads('* [[link][Node 1]]').children[0] >>> node.get_heading() 'Node 1' >>> node.get_heading(format='raw') '[[link][Node 1]]'
- property level[source]¶
Level attribute of this node. Top level node is level 1.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... ''') >>> (n1, n2) = list(root[1:]) >>> root.level 0 >>> n1.level 1 >>> n2.level 2
- property priority: str | None[source]¶
Priority attribute of this node. It is None if undefined.
>>> from orgparse import loads >>> (n1, n2) = loads(''' ... * [#A] Node 1 ... * Node 2 ... ''').children >>> n1.priority 'A' >>> n2.priority is None True
- property todo: str | None[source]¶
A TODO keyword of this node if exists or None otherwise.
>>> from orgparse import loads >>> root = loads('* TODO Node 1') >>> root.children[0].todo 'TODO'
- property scheduled[source]¶
Return scheduled timestamp
- Return type:
a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... SCHEDULED: <2012-02-26 Sun> ... ''') >>> root.children[0].scheduled OrgDateScheduled((2012, 2, 26))
- property deadline[source]¶
Return deadline timestamp.
- Return type:
a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... DEADLINE: <2012-02-26 Sun> ... ''') >>> root.children[0].deadline OrgDateDeadline((2012, 2, 26))
- property closed[source]¶
Return timestamp of closed time.
- Return type:
a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... CLOSED: [2012-02-26 Sun 21:15] ... ''') >>> root.children[0].closed OrgDateClosed((2012, 2, 26, 21, 15, 0))
- property clock[source]¶
Return a list of clocked timestamps
- Return type:
a list of a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05 ... ''') >>> root.children[0].clock [OrgDateClock((2012, 2, 26, 21, 10, 0), (2012, 2, 26, 21, 15, 0))]
- property repeated_tasks[source]¶
Get repeated tasks marked DONE in an entry having repeater.
- Return type:
>>> from orgparse import loads >>> node = loads(''' ... * TODO Pay the rent ... DEADLINE: <2005-10-01 Sat +1m> ... - State "DONE" from "TODO" [2005-09-01 Thu 16:10] ... - State "DONE" from "TODO" [2005-08-01 Mon 19:44] ... - State "DONE" from "TODO" [2005-07-01 Fri 17:27] ... ''').children[0] >>> node.repeated_tasks [OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')] >>> node.repeated_tasks[0].before 'TODO' >>> node.repeated_tasks[0].after 'DONE'
Repeated tasks in
:LOGBOOK:can be fetched by the same code.>>> node = loads(''' ... * TODO Pay the rent ... DEADLINE: <2005-10-01 Sat +1m> ... :LOGBOOK: ... - State "DONE" from "TODO" [2005-09-01 Thu 16:10] ... - State "DONE" from "TODO" [2005-08-01 Mon 19:44] ... - State "DONE" from "TODO" [2005-07-01 Fri 17:27] ... :END: ... ''').children[0] >>> node.repeated_tasks [OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')]
- class orgparse.OrgRootNode(env: OrgEnv, index: int | None = None)[source]¶
Node to represent a file. Its body contains all lines before the first headline
See
OrgBaseNodefor other available functions.- get_parent(max_level=None)[source]¶
Return a parent node.
- Parameters:
max_level (int) –
In the normally structured org file, it is a level of the ancestor node to return. For example,
get_parent(max_level=0)returns a root node.In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to
max_level, this function try to find an ancestor node which level is smaller thanmax_level.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
For simplicity, accessing
parentis alias of callingget_parent()without argument.>>> n1.get_parent() is n1.parent True >>> root.parent is None True
This is a little bit pathological situation – but works.
>>> root = loads(''' ... * Node 1 ... *** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
Now let’s play with max_level.
>>> root = loads(''' ... * Node 1 (level 1) ... ** Node 2 (level 2) ... *** Node 3 (level 3) ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n3.get_parent() is n2 True >>> n3.get_parent(max_level=2) is n2 # same as default True >>> n3.get_parent(max_level=1) is n1 True >>> n3.get_parent(max_level=0) is root True
- orgparse.load(path: str | Path | TextIO, env: OrgEnv | None = None) OrgRootNode[source]¶
Load org-mode document from a file.
- Parameters:
path (str or file-like) – Path to org file or file-like object of an org document.
- Return type:
- orgparse.loadi(lines: Iterable[str], filename: str = '<lines>', env: OrgEnv | None = None) OrgRootNode[source]¶
Load org-mode document from an iterative object.
- Return type:
Tree structure interface¶
- class orgparse.node.OrgBaseNode(env: OrgEnv, index: int | None = None)[source]¶
Base class for
OrgRootNodeandOrgNode- env
An instance of
OrgEnv. All nodes in a same file shares same instance.
OrgBaseNodeis an iterable object:>>> from orgparse import loads >>> root = loads(''' ... * Heading 1 ... ** Heading 2 ... *** Heading 3 ... * Heading 4 ... ''') >>> for node in root: ... print(node) * Heading 1 ** Heading 2 *** Heading 3 * Heading 4
Note that the first blank line is due to the root node, as iteration contains the object itself. To skip that, use slice access
[1:]:>>> for node in root[1:]: ... print(node) * Heading 1 ** Heading 2 *** Heading 3 * Heading 4
It also supports sequence protocol.
>>> print(root[1]) * Heading 1 >>> root[0] is root # index 0 means itself True >>> len(root) # remember, sequence contains itself 5
Note the difference between
root[1:]androot[1]:>>> for node in root[1]: ... print(node) * Heading 1 ** Heading 2 *** Heading 3
Nodes remember the line number information (1-indexed):
>>> print(root.children[1].linenumber) 5
- property end_linenumber: int[source]¶
One-based, inclusive end line of this node’s own source text.
Includes the heading, metadata, and trailing blank lines, excluding descendants. For
OrgRootNode, this covers the preamble before the first heading and returns 0 when it is empty. Usenode[-1].end_linenumberfor the end of the entire subtree.>>> from orgparse import loads >>> root = loads('''Preamble ... * Parent ... body ... ** Child ... child body ... * Last''') >>> [(node.linenumber, node.end_linenumber) for node in root] [(1, 1), (2, 3), (4, 5), (6, 6)] >>> root[1][-1].end_linenumber 5 >>> loads('* Heading').end_linenumber 0
- property previous_same_level: OrgBaseNode | None[source]¶
Return previous node if exists or None otherwise.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... * Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.previous_same_level is None True >>> n2.previous_same_level is n1 True >>> n3.previous_same_level is None # n2 is not at the same level True
- property next_same_level: OrgBaseNode | None[source]¶
Return next node if exists or None otherwise.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... * Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.next_same_level is n2 True >>> n2.next_same_level is None # n3 is not at the same level True >>> n3.next_same_level is None True
- get_parent(max_level: int | None = None)[source]¶
Return a parent node.
- Parameters:
max_level (int) –
In the normally structured org file, it is a level of the ancestor node to return. For example,
get_parent(max_level=0)returns a root node.In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to
max_level, this function try to find an ancestor node which level is smaller thanmax_level.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
For simplicity, accessing
parentis alias of callingget_parent()without argument.>>> n1.get_parent() is n1.parent True >>> root.parent is None True
This is a little bit pathological situation – but works.
>>> root = loads(''' ... * Node 1 ... *** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
Now let’s play with max_level.
>>> root = loads(''' ... * Node 1 (level 1) ... ** Node 2 (level 2) ... *** Node 3 (level 3) ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n3.get_parent() is n2 True >>> n3.get_parent(max_level=2) is n2 # same as default True >>> n3.get_parent(max_level=1) is n1 True >>> n3.get_parent(max_level=0) is root True
- property parent[source]¶
Alias of
get_parent()(calling without argument).
- property children[source]¶
A list of child nodes.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... *** Node 3 ... ** Node 4 ... ''') >>> (n1, n2, n3, n4) = list(root[1:]) >>> (c1, c2) = n1.children >>> c1 is n2 True >>> c2 is n4 True
Note the difference to
n1[1:], which returns the Node 3 also:>>> (m1, m2, m3) = list(n1[1:]) >>> m2 is n3 True
- property root[source]¶
The root node.
>>> from orgparse import loads >>> root = loads('* Node 1') >>> n1 = root[1] >>> n1.root is root True
- property properties: dict[str, str | int | float][source]¶
Node properties as a dictionary.
>>> from orgparse import loads >>> root = loads(''' ... * Node ... :PROPERTIES: ... :SomeProperty: value ... :END: ... ''') >>> root.children[0].properties['SomeProperty'] 'value'
- get_property(key, val=None) str | int | float | None[source]¶
Return property named
keyif exists orvalotherwise.- Parameters:
key (str) – Key of property.
val – Default value to return.
- property tags: set[str][source]¶
Tags of this and parent’s node.
>>> from orgparse import loads >>> n2 = loads(''' ... * Node 1 :TAG1: ... ** Node 2 :TAG2: ... ''')[2] >>> n2.tags == set(['TAG1', 'TAG2']) True
- property shallow_tags: set[str][source]¶
Tags defined for this node (don’t look-up parent nodes).
>>> from orgparse import loads >>> n2 = loads(''' ... * Node 1 :TAG1: ... ** Node 2 :TAG2: ... ''')[2] >>> n2.shallow_tags == set(['TAG2']) True
- is_root()[source]¶
Return
Truewhen it is a root node.>>> from orgparse import loads >>> root = loads('* Node 1') >>> root.is_root() True >>> n1 = root[1] >>> n1.is_root() False
- get_timestamps(active=False, inactive=False, range=False, point=False)[source]¶
Return a list of timestamps in the body text.
- Parameters:
- Return type:
list of
orgparse.date.OrgDatesubclasses
Consider the following org node:
>>> from orgparse import loads >>> node = loads(''' ... * Node ... CLOSED: [2012-02-26 Sun 21:15] SCHEDULED: <2012-02-26 Sun> ... CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05 ... Some inactive timestamp [2012-02-23 Thu] in body text. ... Some active timestamp <2012-02-24 Fri> in body text. ... Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon]. ... Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>. ... ''').children[0]
The default flags are all off, so it does not return anything.
>>> node.get_timestamps() []
You can fetch appropriate timestamps using keyword arguments.
>>> node.get_timestamps(inactive=True, point=True) [OrgDate((2012, 2, 23), None, False)] >>> node.get_timestamps(active=True, point=True) [OrgDate((2012, 2, 24))] >>> node.get_timestamps(inactive=True, range=True) [OrgDate((2012, 2, 25), (2012, 2, 27), False)] >>> node.get_timestamps(active=True, range=True) [OrgDate((2012, 2, 26), (2012, 2, 28))]
This is more complex example. Only active timestamps, regardless of range/point type.
>>> node.get_timestamps(active=True, point=True, range=True) [OrgDate((2012, 2, 24)), OrgDate((2012, 2, 26), (2012, 2, 28))]
- property datelist[source]¶
Alias of
.get_timestamps(active=True, inactive=True, point=True).- Return type:
list of
orgparse.date.OrgDatesubclasses
>>> from orgparse import loads >>> root = loads(''' ... * Node with point dates <2012-02-25 Sat> ... CLOSED: [2012-02-25 Sat 21:15] ... Some inactive timestamp [2012-02-26 Sun] in body text. ... Some active timestamp <2012-02-27 Mon> in body text. ... ''') >>> root.children[0].datelist [OrgDate((2012, 2, 25)), OrgDate((2012, 2, 26), None, False), OrgDate((2012, 2, 27))]
- property rangelist[source]¶
Alias of
.get_timestamps(active=True, inactive=True, range=True).- Return type:
list of
orgparse.date.OrgDatesubclasses
>>> from orgparse import loads >>> root = loads(''' ... * Node with range dates <2012-02-25 Sat>--<2012-02-28 Tue> ... CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05 ... Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon]. ... Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>. ... Some time interval <2012-02-27 Mon 11:23-12:10>. ... ''') >>> root.children[0].rangelist [OrgDate((2012, 2, 25), (2012, 2, 28)), OrgDate((2012, 2, 25), (2012, 2, 27), False), OrgDate((2012, 2, 26), (2012, 2, 28)), OrgDate((2012, 2, 27, 11, 23, 0), (2012, 2, 27, 12, 10, 0))]
- class orgparse.node.OrgRootNode(env: OrgEnv, index: int | None = None)[source]¶
Node to represent a file. Its body contains all lines before the first headline
See
OrgBaseNodefor other available functions.- get_parent(max_level=None)[source]¶
Return a parent node.
- Parameters:
max_level (int) –
In the normally structured org file, it is a level of the ancestor node to return. For example,
get_parent(max_level=0)returns a root node.In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to
max_level, this function try to find an ancestor node which level is smaller thanmax_level.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
For simplicity, accessing
parentis alias of callingget_parent()without argument.>>> n1.get_parent() is n1.parent True >>> root.parent is None True
This is a little bit pathological situation – but works.
>>> root = loads(''' ... * Node 1 ... *** Node 2 ... ** Node 3 ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n1.get_parent() is root True >>> n2.get_parent() is n1 True >>> n3.get_parent() is n1 True
Now let’s play with max_level.
>>> root = loads(''' ... * Node 1 (level 1) ... ** Node 2 (level 2) ... *** Node 3 (level 3) ... ''') >>> (n1, n2, n3) = list(root[1:]) >>> n3.get_parent() is n2 True >>> n3.get_parent(max_level=2) is n2 # same as default True >>> n3.get_parent(max_level=1) is n1 True >>> n3.get_parent(max_level=0) is root True
- class orgparse.node.OrgNode(*args, **kwds)[source]¶
Node to represent normal org node
See
OrgBaseNodefor other available functions.- get_heading(format: str = 'plain') str[source]¶
Return a string of head text without tags and TODO keywords.
>>> from orgparse import loads >>> node = loads('* TODO Node 1').children[0] >>> node.get_heading() 'Node 1'
It strips off inline markup by default (
format='plain'). You can get the original raw string by specifyingformat='raw'.>>> node = loads('* [[link][Node 1]]').children[0] >>> node.get_heading() 'Node 1' >>> node.get_heading(format='raw') '[[link][Node 1]]'
- property level[source]¶
Level attribute of this node. Top level node is level 1.
>>> from orgparse import loads >>> root = loads(''' ... * Node 1 ... ** Node 2 ... ''') >>> (n1, n2) = list(root[1:]) >>> root.level 0 >>> n1.level 1 >>> n2.level 2
- property priority: str | None[source]¶
Priority attribute of this node. It is None if undefined.
>>> from orgparse import loads >>> (n1, n2) = loads(''' ... * [#A] Node 1 ... * Node 2 ... ''').children >>> n1.priority 'A' >>> n2.priority is None True
- property todo: str | None[source]¶
A TODO keyword of this node if exists or None otherwise.
>>> from orgparse import loads >>> root = loads('* TODO Node 1') >>> root.children[0].todo 'TODO'
- property scheduled[source]¶
Return scheduled timestamp
- Return type:
a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... SCHEDULED: <2012-02-26 Sun> ... ''') >>> root.children[0].scheduled OrgDateScheduled((2012, 2, 26))
- property deadline[source]¶
Return deadline timestamp.
- Return type:
a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... DEADLINE: <2012-02-26 Sun> ... ''') >>> root.children[0].deadline OrgDateDeadline((2012, 2, 26))
- property closed[source]¶
Return timestamp of closed time.
- Return type:
a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... CLOSED: [2012-02-26 Sun 21:15] ... ''') >>> root.children[0].closed OrgDateClosed((2012, 2, 26, 21, 15, 0))
- property clock[source]¶
Return a list of clocked timestamps
- Return type:
a list of a subclass of
orgparse.date.OrgDate
>>> from orgparse import loads >>> root = loads(''' ... * Node ... CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05 ... ''') >>> root.children[0].clock [OrgDateClock((2012, 2, 26, 21, 10, 0), (2012, 2, 26, 21, 15, 0))]
- property repeated_tasks[source]¶
Get repeated tasks marked DONE in an entry having repeater.
- Return type:
>>> from orgparse import loads >>> node = loads(''' ... * TODO Pay the rent ... DEADLINE: <2005-10-01 Sat +1m> ... - State "DONE" from "TODO" [2005-09-01 Thu 16:10] ... - State "DONE" from "TODO" [2005-08-01 Mon 19:44] ... - State "DONE" from "TODO" [2005-07-01 Fri 17:27] ... ''').children[0] >>> node.repeated_tasks [OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')] >>> node.repeated_tasks[0].before 'TODO' >>> node.repeated_tasks[0].after 'DONE'
Repeated tasks in
:LOGBOOK:can be fetched by the same code.>>> node = loads(''' ... * TODO Pay the rent ... DEADLINE: <2005-10-01 Sat +1m> ... :LOGBOOK: ... - State "DONE" from "TODO" [2005-09-01 Thu 16:10] ... - State "DONE" from "TODO" [2005-08-01 Mon 19:44] ... - State "DONE" from "TODO" [2005-07-01 Fri 17:27] ... :END: ... ''').children[0] >>> node.repeated_tasks [OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'), OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')]
- class orgparse.node.OrgEnv(todos: Sequence[str] | None = None, dones: Sequence[str] | None = None, filename: str | Path = '<undefined>')[source]¶
Information global to the file (e.g, TODO keywords).
- property nodes: list[OrgBaseNode][source]¶
A list of org nodes.
>>> OrgEnv().nodes # default is empty (of course) []
>>> from orgparse import loads >>> loads(''' ... * Heading 1 ... ** Heading 2 ... *** Heading 3 ... ''').env.nodes [<orgparse.node.OrgRootNode object at 0x...>, <orgparse.node.OrgNode object at 0x...>, <orgparse.node.OrgNode object at 0x...>, <orgparse.node.OrgNode object at 0x...>]
- property todo_keys[source]¶
TODO keywords defined for this document (file).
>>> env = OrgEnv() >>> env.todo_keys ['TODO']
- property done_keys[source]¶
DONE keywords defined for this document (file).
>>> env = OrgEnv() >>> env.done_keys ['DONE']
- property all_todo_keys[source]¶
All TODO keywords (including DONEs).
>>> env = OrgEnv() >>> env.all_todo_keys ['TODO', 'DONE']
- property filename: str[source]¶
Return the source filename as a string.
A
pathlib.Pathpassed toOrgEnvis converted to a string. Documents loaded without a filename use a placeholder such as<string>.
Date interface¶
- class orgparse.date.OrgDate(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]¶
- __init__(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None) None[source]¶
Create
OrgDateobject- Parameters:
start (datetime, date, tuple, int, float or None) – Starting date.
end (datetime, date, tuple, int, float or None) – Ending date.
active (bool or None) – Active/inactive flag. None means using its default value, which may be different for different subclasses.
repeater (tuple or None) – Repeater interval.
warning (tuple or None) – Deadline warning interval.
>>> OrgDate(datetime.date(2012, 2, 10)) OrgDate((2012, 2, 10)) >>> OrgDate((2012, 2, 10)) OrgDate((2012, 2, 10)) >>> OrgDate((2012, 2)) Traceback (most recent call last): ... ValueError: Automatic conversion to the datetime object requires at least 3 elements in the tuple. Only 2 elements are in the given tuple '(2012, 2)'. >>> OrgDate((2012, 2, 10, 12, 20, 30)) OrgDate((2012, 2, 10, 12, 20, 30)) >>> OrgDate((2012, 2, 10), (2012, 2, 15), active=False) OrgDate((2012, 2, 10), (2012, 2, 15), False)
OrgDate can be created using unix timestamp:
>>> OrgDate(datetime.datetime.fromtimestamp(0)) == OrgDate(0) True
- property start: date | datetime[source]¶
Get date or datetime object
>>> OrgDate((2012, 2, 10)).start datetime.date(2012, 2, 10) >>> OrgDate((2012, 2, 10, 12, 10)).start datetime.datetime(2012, 2, 10, 12, 10)
- property end: date | datetime[source]¶
Get date or datetime object
>>> OrgDate((2012, 2, 10), (2012, 2, 15)).end datetime.date(2012, 2, 15) >>> OrgDate((2012, 2, 10, 12, 10), (2012, 2, 15, 12, 10)).end datetime.datetime(2012, 2, 15, 12, 10)
- has_time() bool[source]¶
Return true if the start date has time field
>>> OrgDate((2012, 2, 10)).has_time() False >>> OrgDate((2012, 2, 10, 12, 10)).has_time() True
- has_overlap(other) bool[source]¶
Test if it has overlap with other
OrgDateinstanceIf the argument is not an instance of
OrgDate, it is converted toOrgDateinstance byOrgDate(other)first.>>> od = OrgDate((2012, 2, 10), (2012, 2, 15)) >>> od.has_overlap(OrgDate((2012, 2, 11))) True >>> od.has_overlap(OrgDate((2012, 2, 20))) False >>> od.has_overlap(OrgDate((2012, 2, 11), (2012, 2, 20))) True >>> od.has_overlap((2012, 2, 11)) True
- classmethod list_from_str(string: str) list[OrgDate][source]¶
Parse string and return a list of
OrgDateobjects>>> OrgDate.list_from_str("... <2012-02-10 Fri> and <2012-02-12 Sun>") [OrgDate((2012, 2, 10)), OrgDate((2012, 2, 12))] >>> OrgDate.list_from_str("<2012-02-10 Fri>--<2012-02-12 Sun>") [OrgDate((2012, 2, 10), (2012, 2, 12))] >>> OrgDate.list_from_str("<2012-02-10 Fri>--[2012-02-12 Sun]") [OrgDate((2012, 2, 10)), OrgDate((2012, 2, 12), None, False)] >>> OrgDate.list_from_str("this is not timestamp") [] >>> OrgDate.list_from_str("<2012-02-11 Sat 10:11--11:20>") [OrgDate((2012, 2, 11, 10, 11, 0), (2012, 2, 11, 11, 20, 0))]
- class orgparse.date.OrgDateScheduled(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]¶
Date object to represent SCHEDULED attribute.
- class orgparse.date.OrgDateDeadline(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]¶
Date object to represent DEADLINE attribute.
- class orgparse.date.OrgDateClosed(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]¶
Date object to represent CLOSED attribute.
- class orgparse.date.OrgDateClock(start, end=None, duration=None, active=None)[source]¶
Date object to represent CLOCK attributes.
>>> OrgDateClock.from_str( ... 'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] => 0:30') OrgDateClock((2010, 8, 8, 17, 0, 0), (2010, 8, 8, 17, 30, 0))
- property duration[source]¶
Get duration of CLOCK.
>>> duration = OrgDateClock.from_str( ... 'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] => 0:30' ... ).duration >>> duration.seconds 1800 >>> total_minutes(duration) 30.0
- is_duration_consistent()[source]¶
Check duration value of CLOCK line.
>>> OrgDateClock.from_str( ... 'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] => 0:30' ... ).is_duration_consistent() True >>> OrgDateClock.from_str( ... 'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] => 0:15' ... ).is_duration_consistent() False
- classmethod from_str(string: str) OrgDateClock[source]¶
Get CLOCK from given string.
Return three tuple (start, stop, length) which is datetime object of start time, datetime object of stop time and length in minute.