orgparse — Python module for reading Emacs org-mode files

Install

You can install orgparse via PyPI:

pip install orgparse

Or via conda-forge:

conda install orgparse -c conda-forge

Usage

The API documentation includes extensive doctests for individual methods. Here are some examples to get started.

Load an org document

from orgparse import load, loads

load('PATH/TO/FILE.org')
load(file_like_object)

loads('''
* This is org-mode contents
  You can load org object from string.
** Second header
''')

See the loading implementation.

Traverse an org tree

>>> from orgparse import loads
>>> root = loads('''
... * Heading 1
... ** Heading 2
... *** Heading 3
... ''')
>>> for node in root[1:]:  # [1:] for skipping root itself
...     print(node)
* Heading 1
** Heading 2
*** Heading 3
>>> h1 = root.children[0]
>>> h2 = h1.children[0]
>>> h3 = h2.children[0]
>>> print(h1)
* Heading 1
>>> print(h2)
** Heading 2
>>> print(h3)
*** Heading 3
>>> print(h2.get_parent())
* Heading 1
>>> print(h3.get_parent(max_level=1))
* Heading 1

Access node attributes

>>> root = loads('''
... * DONE Heading          :TAG:
...   CLOSED: [2012-02-26 Sun 21:15] SCHEDULED: <2012-02-26 Sun>
...   CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] =>  0:05
...   :PROPERTIES:
...   :Effort:   1:00
...   :OtherProperty:   some text
...   :END:
...   Body texts...
... ''')
>>> node = root.children[0]
>>> node.heading
'Heading'
>>> node.scheduled
OrgDateScheduled((2012, 2, 26))
>>> node.closed
OrgDateClosed((2012, 2, 26, 21, 15, 0))
>>> node.clock
[OrgDateClock((2012, 2, 26, 21, 10, 0), (2012, 2, 26, 21, 15, 0))]
>>> bool(node.deadline)  # it is not specified
False
>>> node.tags == set(['TAG'])
True
>>> node.get_property('Effort')
60
>>> node.get_property('UndefinedProperty')  # returns None
>>> node.get_property('OtherProperty')
'some text'
>>> node.body
'  Body texts...'

Read named tables

Tables in node.body_rich expose their #+NAME: through Table.name. The name is None for unnamed tables.

>>> from orgparse.extra import Table
>>> root = loads('''
... #+NAME: measurements
... | x | y |
... |---+---|
... | 1 | 2 |
... ''')
>>> [table] = [part for part in root.body_rich if isinstance(part, Table) and part.name == 'measurements']
>>> list(table.as_dicts)
[{'x': '1', 'y': '2'}]

More examples

The tests show additional supported features:

Development and documentation

Clone with git clone --recurse-submodules https://github.com/karlicoss/orgparse.git to include the test corpus. For an existing checkout, run git submodule update --init --recursive.

Run the tests with uv tool run --with tox-uv tox -e tests. This also checks the examples in this README.

Corpus tests parse upstream Org examples, Sacha Chua’s Emacs configuration, exobrain notes, and nvim-orgmode documents from testdata/external. They check tree structure, source line ranges, and attribute access.

Edit README.qmd, then regenerate README.md with uv tool run --with tox-uv tox -e quarto. Quarto computes links to source code and tests from their definitions, so line numbers are refreshed when rendering. Commit both files together; CI checks that the generated README is current.

Build the documentation with uv tool run --with tox-uv tox -e docs and open doc/_build/html/index.html. Sphinx combines the generated README with the API docstrings; it does not need Quarto to build the site.

Read the Docs uses .readthedocs.yaml to run the same docs environment. The latest version follows master, and stable follows releases. Automatic builds require the GitHub integration in the Read the Docs project settings.

Project status

The project is maintained by @karlicoss.

For my personal use, orgparse mostly has all features I need, so there hasn’t been much active development lately.

However, contributions are always welcome! Please provide tests along with your contribution if you’re fixing bugs or adding new functionality.

Loading documents

Read Emacs org-mode files as trees of Python objects.

class orgparse.OrgBaseNode(env: OrgEnv, index: int | None = None)[source]

Base class for OrgRootNode and OrgNode

env

An instance of OrgEnv. All nodes in a same file shares same instance.

OrgBaseNode is an iterable object:

>>> from orgparse import loads
>>> root = loads('''
... * Heading 1
... ** Heading 2
... *** Heading 3
... * Heading 4
... ''')
>>> for node in root:
...     print(node)

* Heading 1
** Heading 2
*** Heading 3
* Heading 4

Note that the first blank line is due to the root node, as iteration contains the object itself. To skip that, use slice access [1:]:

>>> for node in root[1:]:
...     print(node)
* Heading 1
** Heading 2
*** Heading 3
* Heading 4

It also supports sequence protocol.

>>> print(root[1])
* Heading 1
>>> root[0] is root  # index 0 means itself
True
>>> len(root)   # remember, sequence contains itself
5

Note the difference between root[1:] and root[1]:

>>> for node in root[1]:
...     print(node)
* Heading 1
** Heading 2
*** Heading 3

Nodes remember the line number information (1-indexed):

>>> print(root.children[1].linenumber)
5
property end_linenumber: int[source]

One-based, inclusive end line of this node’s own source text.

Includes the heading, metadata, and trailing blank lines, excluding descendants. For OrgRootNode, this covers the preamble before the first heading and returns 0 when it is empty. Use node[-1].end_linenumber for the end of the entire subtree.

>>> from orgparse import loads
>>> root = loads('''Preamble
... * Parent
... body
... ** Child
... child body
... * Last''')
>>> [(node.linenumber, node.end_linenumber) for node in root]
[(1, 1), (2, 3), (4, 5), (6, 6)]
>>> root[1][-1].end_linenumber
5
>>> loads('* Heading').end_linenumber
0
property previous_same_level: OrgBaseNode | None[source]

Return previous node if exists or None otherwise.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... * Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.previous_same_level is None
True
>>> n2.previous_same_level is n1
True
>>> n3.previous_same_level is None  # n2 is not at the same level
True
property next_same_level: OrgBaseNode | None[source]

Return next node if exists or None otherwise.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... * Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.next_same_level is n2
True
>>> n2.next_same_level is None  # n3 is not at the same level
True
>>> n3.next_same_level is None
True
get_parent(max_level: int | None = None)[source]

Return a parent node.

Parameters:

max_level (int) –

In the normally structured org file, it is a level of the ancestor node to return. For example, get_parent(max_level=0) returns a root node.

In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to max_level, this function try to find an ancestor node which level is smaller than max_level.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

For simplicity, accessing parent is alias of calling get_parent() without argument.

>>> n1.get_parent() is n1.parent
True
>>> root.parent is None
True

This is a little bit pathological situation – but works.

>>> root = loads('''
... * Node 1
... *** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

Now let’s play with max_level.

>>> root = loads('''
... * Node 1 (level 1)
... ** Node 2 (level 2)
... *** Node 3 (level 3)
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n3.get_parent() is n2
True
>>> n3.get_parent(max_level=2) is n2  # same as default
True
>>> n3.get_parent(max_level=1) is n1
True
>>> n3.get_parent(max_level=0) is root
True
property parent[source]

Alias of get_parent() (calling without argument).

property children[source]

A list of child nodes.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... *** Node 3
... ** Node 4
... ''')
>>> (n1, n2, n3, n4) = list(root[1:])
>>> (c1, c2) = n1.children
>>> c1 is n2
True
>>> c2 is n4
True

Note the difference to n1[1:], which returns the Node 3 also:

>>> (m1, m2, m3) = list(n1[1:])
>>> m2 is n3
True
property root[source]

The root node.

>>> from orgparse import loads
>>> root = loads('* Node 1')
>>> n1 = root[1]
>>> n1.root is root
True
property properties: dict[str, str | int | float][source]

Node properties as a dictionary.

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   :PROPERTIES:
...   :SomeProperty: value
...   :END:
... ''')
>>> root.children[0].properties['SomeProperty']
'value'
get_property(key, val=None) → str | int | float | None[source]

Return property named key if exists or val otherwise.

Parameters:
  • key (str) – Key of property.

  • val – Default value to return.

property level: int[source]

Level of this node.

property tags: set[str][source]

Tags of this and parent’s node.

>>> from orgparse import loads
>>> n2 = loads('''
... * Node 1    :TAG1:
... ** Node 2   :TAG2:
... ''')[2]
>>> n2.tags == set(['TAG1', 'TAG2'])
True
property shallow_tags: set[str][source]

Tags defined for this node (don’t look-up parent nodes).

>>> from orgparse import loads
>>> n2 = loads('''
... * Node 1    :TAG1:
... ** Node 2   :TAG2:
... ''')[2]
>>> n2.shallow_tags == set(['TAG2'])
True
get_body(format: str = 'plain') → str[source]

Return a string of body text.

See also: get_heading().

property body: str[source]

Alias of .get_body(format='plain').

is_root()[source]

Return True when it is a root node.

>>> from orgparse import loads
>>> root = loads('* Node 1')
>>> root.is_root()
True
>>> n1 = root[1]
>>> n1.is_root()
False
get_timestamps(active=False, inactive=False, range=False, point=False)[source]

Return a list of timestamps in the body text.

Parameters:
  • active (bool) – Include active type timestamps.

  • inactive (bool) – Include inactive type timestamps.

  • range (bool) – Include timestamps which has end date.

  • point (bool) – Include timestamps which has no end date.

Return type:

list of orgparse.date.OrgDate subclasses

Consider the following org node:

>>> from orgparse import loads
>>> node = loads('''
... * Node
...   CLOSED: [2012-02-26 Sun 21:15] SCHEDULED: <2012-02-26 Sun>
...   CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] =>  0:05
...   Some inactive timestamp [2012-02-23 Thu] in body text.
...   Some active timestamp <2012-02-24 Fri> in body text.
...   Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon].
...   Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>.
... ''').children[0]

The default flags are all off, so it does not return anything.

>>> node.get_timestamps()
[]

You can fetch appropriate timestamps using keyword arguments.

>>> node.get_timestamps(inactive=True, point=True)
[OrgDate((2012, 2, 23), None, False)]
>>> node.get_timestamps(active=True, point=True)
[OrgDate((2012, 2, 24))]
>>> node.get_timestamps(inactive=True, range=True)
[OrgDate((2012, 2, 25), (2012, 2, 27), False)]
>>> node.get_timestamps(active=True, range=True)
[OrgDate((2012, 2, 26), (2012, 2, 28))]

This is more complex example. Only active timestamps, regardless of range/point type.

>>> node.get_timestamps(active=True, point=True, range=True)
[OrgDate((2012, 2, 24)), OrgDate((2012, 2, 26), (2012, 2, 28))]
property datelist[source]

Alias of .get_timestamps(active=True, inactive=True, point=True).

Return type:

list of orgparse.date.OrgDate subclasses

>>> from orgparse import loads
>>> root = loads('''
... * Node with point dates <2012-02-25 Sat>
...   CLOSED: [2012-02-25 Sat 21:15]
...   Some inactive timestamp [2012-02-26 Sun] in body text.
...   Some active timestamp <2012-02-27 Mon> in body text.
... ''')
>>> root.children[0].datelist
[OrgDate((2012, 2, 25)),
 OrgDate((2012, 2, 26), None, False),
 OrgDate((2012, 2, 27))]
property rangelist[source]

Alias of .get_timestamps(active=True, inactive=True, range=True).

Return type:

list of orgparse.date.OrgDate subclasses

>>> from orgparse import loads
>>> root = loads('''
... * Node with range dates <2012-02-25 Sat>--<2012-02-28 Tue>
...   CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05
...   Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon].
...   Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>.
...   Some time interval <2012-02-27 Mon 11:23-12:10>.
... ''')
>>> root.children[0].rangelist
[OrgDate((2012, 2, 25), (2012, 2, 28)),
 OrgDate((2012, 2, 25), (2012, 2, 27), False),
 OrgDate((2012, 2, 26), (2012, 2, 28)),
 OrgDate((2012, 2, 27, 11, 23, 0), (2012, 2, 27, 12, 10, 0))]
get_file_property_list(property: str)[source]

Return a list of the selected property

get_file_property(property: str)[source]

Return a single element of the selected property or None if it doesn’t exist

class orgparse.OrgEnv(todos: Sequence[str] | None = None, dones: Sequence[str] | None = None, filename: str | Path = '<undefined>')[source]

Information global to the file (e.g, TODO keywords).

property nodes: list[OrgBaseNode][source]

A list of org nodes.

>>> OrgEnv().nodes   # default is empty (of course)
[]
>>> from orgparse import loads
>>> loads('''
... * Heading 1
... ** Heading 2
... *** Heading 3
... ''').env.nodes
[<orgparse.node.OrgRootNode object at 0x...>,
 <orgparse.node.OrgNode object at 0x...>,
 <orgparse.node.OrgNode object at 0x...>,
 <orgparse.node.OrgNode object at 0x...>]
property todo_keys[source]

TODO keywords defined for this document (file).

>>> env = OrgEnv()
>>> env.todo_keys
['TODO']
property done_keys[source]

DONE keywords defined for this document (file).

>>> env = OrgEnv()
>>> env.done_keys
['DONE']
property all_todo_keys[source]

All TODO keywords (including DONEs).

>>> env = OrgEnv()
>>> env.all_todo_keys
['TODO', 'DONE']
property filename: str[source]

Return the source filename as a string.

A pathlib.Path passed to OrgEnv is converted to a string. Documents loaded without a filename use a placeholder such as <string>.

class orgparse.OrgNode(*args, **kwds)[source]

Node to represent normal org node

See OrgBaseNode for other available functions.

get_heading(format: str = 'plain') → str[source]

Return a string of head text without tags and TODO keywords.

>>> from orgparse import loads
>>> node = loads('* TODO Node 1').children[0]
>>> node.get_heading()
'Node 1'

It strips off inline markup by default (format='plain'). You can get the original raw string by specifying format='raw'.

>>> node = loads('* [[link][Node 1]]').children[0]
>>> node.get_heading()
'Node 1'
>>> node.get_heading(format='raw')
'[[link][Node 1]]'
property heading: str[source]

Alias of .get_heading(format='plain').

property level[source]

Level attribute of this node. Top level node is level 1.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... ''')
>>> (n1, n2) = list(root[1:])
>>> root.level
0
>>> n1.level
1
>>> n2.level
2
property priority: str | None[source]

Priority attribute of this node. It is None if undefined.

>>> from orgparse import loads
>>> (n1, n2) = loads('''
... * [#A] Node 1
... * Node 2
... ''').children
>>> n1.priority
'A'
>>> n2.priority is None
True
property todo: str | None[source]

A TODO keyword of this node if exists or None otherwise.

>>> from orgparse import loads
>>> root = loads('* TODO Node 1')
>>> root.children[0].todo
'TODO'
property scheduled[source]

Return scheduled timestamp

Return type:

a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   SCHEDULED: <2012-02-26 Sun>
... ''')
>>> root.children[0].scheduled
OrgDateScheduled((2012, 2, 26))
property deadline[source]

Return deadline timestamp.

Return type:

a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   DEADLINE: <2012-02-26 Sun>
... ''')
>>> root.children[0].deadline
OrgDateDeadline((2012, 2, 26))
property closed[source]

Return timestamp of closed time.

Return type:

a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   CLOSED: [2012-02-26 Sun 21:15]
... ''')
>>> root.children[0].closed
OrgDateClosed((2012, 2, 26, 21, 15, 0))
property clock[source]

Return a list of clocked timestamps

Return type:

a list of a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] =>  0:05
... ''')
>>> root.children[0].clock
[OrgDateClock((2012, 2, 26, 21, 10, 0), (2012, 2, 26, 21, 15, 0))]
has_date()[source]

Return True if it has any kind of timestamp

property repeated_tasks[source]

Get repeated tasks marked DONE in an entry having repeater.

Return type:

list of orgparse.date.OrgDateRepeatedTask

>>> from orgparse import loads
>>> node = loads('''
... * TODO Pay the rent
...   DEADLINE: <2005-10-01 Sat +1m>
...   - State "DONE"  from "TODO"  [2005-09-01 Thu 16:10]
...   - State "DONE"  from "TODO"  [2005-08-01 Mon 19:44]
...   - State "DONE"  from "TODO"  [2005-07-01 Fri 17:27]
... ''').children[0]
>>> node.repeated_tasks
[OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')]
>>> node.repeated_tasks[0].before
'TODO'
>>> node.repeated_tasks[0].after
'DONE'

Repeated tasks in :LOGBOOK: can be fetched by the same code.

>>> node = loads('''
... * TODO Pay the rent
...   DEADLINE: <2005-10-01 Sat +1m>
...   :LOGBOOK:
...   - State "DONE"  from "TODO"  [2005-09-01 Thu 16:10]
...   - State "DONE"  from "TODO"  [2005-08-01 Mon 19:44]
...   - State "DONE"  from "TODO"  [2005-07-01 Fri 17:27]
...   :END:
... ''').children[0]
>>> node.repeated_tasks
[OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')]

See: (info “(org) Repeated tasks”)

class orgparse.OrgRootNode(env: OrgEnv, index: int | None = None)[source]

Node to represent a file. Its body contains all lines before the first headline

See OrgBaseNode for other available functions.

property level: int[source]

Level of this node.

get_parent(max_level=None)[source]

Return a parent node.

Parameters:

max_level (int) –

In the normally structured org file, it is a level of the ancestor node to return. For example, get_parent(max_level=0) returns a root node.

In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to max_level, this function try to find an ancestor node which level is smaller than max_level.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

For simplicity, accessing parent is alias of calling get_parent() without argument.

>>> n1.get_parent() is n1.parent
True
>>> root.parent is None
True

This is a little bit pathological situation – but works.

>>> root = loads('''
... * Node 1
... *** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

Now let’s play with max_level.

>>> root = loads('''
... * Node 1 (level 1)
... ** Node 2 (level 2)
... *** Node 3 (level 3)
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n3.get_parent() is n2
True
>>> n3.get_parent(max_level=2) is n2  # same as default
True
>>> n3.get_parent(max_level=1) is n1
True
>>> n3.get_parent(max_level=0) is root
True
is_root() → bool[source]

Return True when it is a root node.

>>> from orgparse import loads
>>> root = loads('* Node 1')
>>> root.is_root()
True
>>> n1 = root[1]
>>> n1.is_root()
False
orgparse.load(path: str | Path | TextIO, env: OrgEnv | None = None) → OrgRootNode[source]

Load org-mode document from a file.

Parameters:

path (str or file-like) – Path to org file or file-like object of an org document.

Return type:

orgparse.node.OrgRootNode

orgparse.loadi(lines: Iterable[str], filename: str = '<lines>', env: OrgEnv | None = None) → OrgRootNode[source]

Load org-mode document from an iterative object.

Return type:

orgparse.node.OrgRootNode

orgparse.loads(string: str, filename: str = '<string>', env: OrgEnv | None = None) → OrgRootNode[source]

Load org-mode document from a string.

Return type:

orgparse.node.OrgRootNode

Tree structure interface

class orgparse.node.OrgBaseNode(env: OrgEnv, index: int | None = None)[source]

Base class for OrgRootNode and OrgNode

env

An instance of OrgEnv. All nodes in a same file shares same instance.

OrgBaseNode is an iterable object:

>>> from orgparse import loads
>>> root = loads('''
... * Heading 1
... ** Heading 2
... *** Heading 3
... * Heading 4
... ''')
>>> for node in root:
...     print(node)

* Heading 1
** Heading 2
*** Heading 3
* Heading 4

Note that the first blank line is due to the root node, as iteration contains the object itself. To skip that, use slice access [1:]:

>>> for node in root[1:]:
...     print(node)
* Heading 1
** Heading 2
*** Heading 3
* Heading 4

It also supports sequence protocol.

>>> print(root[1])
* Heading 1
>>> root[0] is root  # index 0 means itself
True
>>> len(root)   # remember, sequence contains itself
5

Note the difference between root[1:] and root[1]:

>>> for node in root[1]:
...     print(node)
* Heading 1
** Heading 2
*** Heading 3

Nodes remember the line number information (1-indexed):

>>> print(root.children[1].linenumber)
5
__init__(env: OrgEnv, index: int | None = None) → None[source]
property end_linenumber: int[source]

One-based, inclusive end line of this node’s own source text.

Includes the heading, metadata, and trailing blank lines, excluding descendants. For OrgRootNode, this covers the preamble before the first heading and returns 0 when it is empty. Use node[-1].end_linenumber for the end of the entire subtree.

>>> from orgparse import loads
>>> root = loads('''Preamble
... * Parent
... body
... ** Child
... child body
... * Last''')
>>> [(node.linenumber, node.end_linenumber) for node in root]
[(1, 1), (2, 3), (4, 5), (6, 6)]
>>> root[1][-1].end_linenumber
5
>>> loads('* Heading').end_linenumber
0
property previous_same_level: OrgBaseNode | None[source]

Return previous node if exists or None otherwise.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... * Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.previous_same_level is None
True
>>> n2.previous_same_level is n1
True
>>> n3.previous_same_level is None  # n2 is not at the same level
True
property next_same_level: OrgBaseNode | None[source]

Return next node if exists or None otherwise.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... * Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.next_same_level is n2
True
>>> n2.next_same_level is None  # n3 is not at the same level
True
>>> n3.next_same_level is None
True
get_parent(max_level: int | None = None)[source]

Return a parent node.

Parameters:

max_level (int) –

In the normally structured org file, it is a level of the ancestor node to return. For example, get_parent(max_level=0) returns a root node.

In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to max_level, this function try to find an ancestor node which level is smaller than max_level.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

For simplicity, accessing parent is alias of calling get_parent() without argument.

>>> n1.get_parent() is n1.parent
True
>>> root.parent is None
True

This is a little bit pathological situation – but works.

>>> root = loads('''
... * Node 1
... *** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

Now let’s play with max_level.

>>> root = loads('''
... * Node 1 (level 1)
... ** Node 2 (level 2)
... *** Node 3 (level 3)
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n3.get_parent() is n2
True
>>> n3.get_parent(max_level=2) is n2  # same as default
True
>>> n3.get_parent(max_level=1) is n1
True
>>> n3.get_parent(max_level=0) is root
True
property parent[source]

Alias of get_parent() (calling without argument).

property children[source]

A list of child nodes.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... *** Node 3
... ** Node 4
... ''')
>>> (n1, n2, n3, n4) = list(root[1:])
>>> (c1, c2) = n1.children
>>> c1 is n2
True
>>> c2 is n4
True

Note the difference to n1[1:], which returns the Node 3 also:

>>> (m1, m2, m3) = list(n1[1:])
>>> m2 is n3
True
property root[source]

The root node.

>>> from orgparse import loads
>>> root = loads('* Node 1')
>>> n1 = root[1]
>>> n1.root is root
True
property properties: dict[str, str | int | float][source]

Node properties as a dictionary.

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   :PROPERTIES:
...   :SomeProperty: value
...   :END:
... ''')
>>> root.children[0].properties['SomeProperty']
'value'
get_property(key, val=None) → str | int | float | None[source]

Return property named key if exists or val otherwise.

Parameters:
  • key (str) – Key of property.

  • val – Default value to return.

property level: int[source]

Level of this node.

property tags: set[str][source]

Tags of this and parent’s node.

>>> from orgparse import loads
>>> n2 = loads('''
... * Node 1    :TAG1:
... ** Node 2   :TAG2:
... ''')[2]
>>> n2.tags == set(['TAG1', 'TAG2'])
True
property shallow_tags: set[str][source]

Tags defined for this node (don’t look-up parent nodes).

>>> from orgparse import loads
>>> n2 = loads('''
... * Node 1    :TAG1:
... ** Node 2   :TAG2:
... ''')[2]
>>> n2.shallow_tags == set(['TAG2'])
True
get_body(format: str = 'plain') → str[source]

Return a string of body text.

See also: get_heading().

property body: str[source]

Alias of .get_body(format='plain').

is_root()[source]

Return True when it is a root node.

>>> from orgparse import loads
>>> root = loads('* Node 1')
>>> root.is_root()
True
>>> n1 = root[1]
>>> n1.is_root()
False
get_timestamps(active=False, inactive=False, range=False, point=False)[source]

Return a list of timestamps in the body text.

Parameters:
  • active (bool) – Include active type timestamps.

  • inactive (bool) – Include inactive type timestamps.

  • range (bool) – Include timestamps which has end date.

  • point (bool) – Include timestamps which has no end date.

Return type:

list of orgparse.date.OrgDate subclasses

Consider the following org node:

>>> from orgparse import loads
>>> node = loads('''
... * Node
...   CLOSED: [2012-02-26 Sun 21:15] SCHEDULED: <2012-02-26 Sun>
...   CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] =>  0:05
...   Some inactive timestamp [2012-02-23 Thu] in body text.
...   Some active timestamp <2012-02-24 Fri> in body text.
...   Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon].
...   Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>.
... ''').children[0]

The default flags are all off, so it does not return anything.

>>> node.get_timestamps()
[]

You can fetch appropriate timestamps using keyword arguments.

>>> node.get_timestamps(inactive=True, point=True)
[OrgDate((2012, 2, 23), None, False)]
>>> node.get_timestamps(active=True, point=True)
[OrgDate((2012, 2, 24))]
>>> node.get_timestamps(inactive=True, range=True)
[OrgDate((2012, 2, 25), (2012, 2, 27), False)]
>>> node.get_timestamps(active=True, range=True)
[OrgDate((2012, 2, 26), (2012, 2, 28))]

This is more complex example. Only active timestamps, regardless of range/point type.

>>> node.get_timestamps(active=True, point=True, range=True)
[OrgDate((2012, 2, 24)), OrgDate((2012, 2, 26), (2012, 2, 28))]
property datelist[source]

Alias of .get_timestamps(active=True, inactive=True, point=True).

Return type:

list of orgparse.date.OrgDate subclasses

>>> from orgparse import loads
>>> root = loads('''
... * Node with point dates <2012-02-25 Sat>
...   CLOSED: [2012-02-25 Sat 21:15]
...   Some inactive timestamp [2012-02-26 Sun] in body text.
...   Some active timestamp <2012-02-27 Mon> in body text.
... ''')
>>> root.children[0].datelist
[OrgDate((2012, 2, 25)),
 OrgDate((2012, 2, 26), None, False),
 OrgDate((2012, 2, 27))]
property rangelist[source]

Alias of .get_timestamps(active=True, inactive=True, range=True).

Return type:

list of orgparse.date.OrgDate subclasses

>>> from orgparse import loads
>>> root = loads('''
... * Node with range dates <2012-02-25 Sat>--<2012-02-28 Tue>
...   CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] => 0:05
...   Some inactive time range [2012-02-25 Sat]--[2012-02-27 Mon].
...   Some active time range <2012-02-26 Sun>--<2012-02-28 Tue>.
...   Some time interval <2012-02-27 Mon 11:23-12:10>.
... ''')
>>> root.children[0].rangelist
[OrgDate((2012, 2, 25), (2012, 2, 28)),
 OrgDate((2012, 2, 25), (2012, 2, 27), False),
 OrgDate((2012, 2, 26), (2012, 2, 28)),
 OrgDate((2012, 2, 27, 11, 23, 0), (2012, 2, 27, 12, 10, 0))]
get_file_property_list(property: str)[source]

Return a list of the selected property

get_file_property(property: str)[source]

Return a single element of the selected property or None if it doesn’t exist

class orgparse.node.OrgRootNode(env: OrgEnv, index: int | None = None)[source]

Node to represent a file. Its body contains all lines before the first headline

See OrgBaseNode for other available functions.

property level: int[source]

Level of this node.

get_parent(max_level=None)[source]

Return a parent node.

Parameters:

max_level (int) –

In the normally structured org file, it is a level of the ancestor node to return. For example, get_parent(max_level=0) returns a root node.

In the general case, it specify a maximum level of the desired ancestor node. If there is no ancestor node whose level is equal to max_level, this function try to find an ancestor node which level is smaller than max_level.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

For simplicity, accessing parent is alias of calling get_parent() without argument.

>>> n1.get_parent() is n1.parent
True
>>> root.parent is None
True

This is a little bit pathological situation – but works.

>>> root = loads('''
... * Node 1
... *** Node 2
... ** Node 3
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n1.get_parent() is root
True
>>> n2.get_parent() is n1
True
>>> n3.get_parent() is n1
True

Now let’s play with max_level.

>>> root = loads('''
... * Node 1 (level 1)
... ** Node 2 (level 2)
... *** Node 3 (level 3)
... ''')
>>> (n1, n2, n3) = list(root[1:])
>>> n3.get_parent() is n2
True
>>> n3.get_parent(max_level=2) is n2  # same as default
True
>>> n3.get_parent(max_level=1) is n1
True
>>> n3.get_parent(max_level=0) is root
True
is_root() → bool[source]

Return True when it is a root node.

>>> from orgparse import loads
>>> root = loads('* Node 1')
>>> root.is_root()
True
>>> n1 = root[1]
>>> n1.is_root()
False
class orgparse.node.OrgNode(*args, **kwds)[source]

Node to represent normal org node

See OrgBaseNode for other available functions.

get_heading(format: str = 'plain') → str[source]

Return a string of head text without tags and TODO keywords.

>>> from orgparse import loads
>>> node = loads('* TODO Node 1').children[0]
>>> node.get_heading()
'Node 1'

It strips off inline markup by default (format='plain'). You can get the original raw string by specifying format='raw'.

>>> node = loads('* [[link][Node 1]]').children[0]
>>> node.get_heading()
'Node 1'
>>> node.get_heading(format='raw')
'[[link][Node 1]]'
property heading: str[source]

Alias of .get_heading(format='plain').

property level[source]

Level attribute of this node. Top level node is level 1.

>>> from orgparse import loads
>>> root = loads('''
... * Node 1
... ** Node 2
... ''')
>>> (n1, n2) = list(root[1:])
>>> root.level
0
>>> n1.level
1
>>> n2.level
2
property priority: str | None[source]

Priority attribute of this node. It is None if undefined.

>>> from orgparse import loads
>>> (n1, n2) = loads('''
... * [#A] Node 1
... * Node 2
... ''').children
>>> n1.priority
'A'
>>> n2.priority is None
True
property todo: str | None[source]

A TODO keyword of this node if exists or None otherwise.

>>> from orgparse import loads
>>> root = loads('* TODO Node 1')
>>> root.children[0].todo
'TODO'
property scheduled[source]

Return scheduled timestamp

Return type:

a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   SCHEDULED: <2012-02-26 Sun>
... ''')
>>> root.children[0].scheduled
OrgDateScheduled((2012, 2, 26))
property deadline[source]

Return deadline timestamp.

Return type:

a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   DEADLINE: <2012-02-26 Sun>
... ''')
>>> root.children[0].deadline
OrgDateDeadline((2012, 2, 26))
property closed[source]

Return timestamp of closed time.

Return type:

a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   CLOSED: [2012-02-26 Sun 21:15]
... ''')
>>> root.children[0].closed
OrgDateClosed((2012, 2, 26, 21, 15, 0))
property clock[source]

Return a list of clocked timestamps

Return type:

a list of a subclass of orgparse.date.OrgDate

>>> from orgparse import loads
>>> root = loads('''
... * Node
...   CLOCK: [2012-02-26 Sun 21:10]--[2012-02-26 Sun 21:15] =>  0:05
... ''')
>>> root.children[0].clock
[OrgDateClock((2012, 2, 26, 21, 10, 0), (2012, 2, 26, 21, 15, 0))]
has_date()[source]

Return True if it has any kind of timestamp

property repeated_tasks[source]

Get repeated tasks marked DONE in an entry having repeater.

Return type:

list of orgparse.date.OrgDateRepeatedTask

>>> from orgparse import loads
>>> node = loads('''
... * TODO Pay the rent
...   DEADLINE: <2005-10-01 Sat +1m>
...   - State "DONE"  from "TODO"  [2005-09-01 Thu 16:10]
...   - State "DONE"  from "TODO"  [2005-08-01 Mon 19:44]
...   - State "DONE"  from "TODO"  [2005-07-01 Fri 17:27]
... ''').children[0]
>>> node.repeated_tasks
[OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')]
>>> node.repeated_tasks[0].before
'TODO'
>>> node.repeated_tasks[0].after
'DONE'

Repeated tasks in :LOGBOOK: can be fetched by the same code.

>>> node = loads('''
... * TODO Pay the rent
...   DEADLINE: <2005-10-01 Sat +1m>
...   :LOGBOOK:
...   - State "DONE"  from "TODO"  [2005-09-01 Thu 16:10]
...   - State "DONE"  from "TODO"  [2005-08-01 Mon 19:44]
...   - State "DONE"  from "TODO"  [2005-07-01 Fri 17:27]
...   :END:
... ''').children[0]
>>> node.repeated_tasks
[OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 8, 1, 19, 44, 0), 'TODO', 'DONE'),
 OrgDateRepeatedTask((2005, 7, 1, 17, 27, 0), 'TODO', 'DONE')]

See: (info “(org) Repeated tasks”)

class orgparse.node.OrgEnv(todos: Sequence[str] | None = None, dones: Sequence[str] | None = None, filename: str | Path = '<undefined>')[source]

Information global to the file (e.g, TODO keywords).

property nodes: list[OrgBaseNode][source]

A list of org nodes.

>>> OrgEnv().nodes   # default is empty (of course)
[]
>>> from orgparse import loads
>>> loads('''
... * Heading 1
... ** Heading 2
... *** Heading 3
... ''').env.nodes
[<orgparse.node.OrgRootNode object at 0x...>,
 <orgparse.node.OrgNode object at 0x...>,
 <orgparse.node.OrgNode object at 0x...>,
 <orgparse.node.OrgNode object at 0x...>]
property todo_keys[source]

TODO keywords defined for this document (file).

>>> env = OrgEnv()
>>> env.todo_keys
['TODO']
property done_keys[source]

DONE keywords defined for this document (file).

>>> env = OrgEnv()
>>> env.done_keys
['DONE']
property all_todo_keys[source]

All TODO keywords (including DONEs).

>>> env = OrgEnv()
>>> env.all_todo_keys
['TODO', 'DONE']
property filename: str[source]

Return the source filename as a string.

A pathlib.Path passed to OrgEnv is converted to a string. Documents loaded without a filename use a placeholder such as <string>.

Date interface

class orgparse.date.OrgDate(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]
__init__(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None) → None[source]

Create OrgDate object

Parameters:
  • start (datetime, date, tuple, int, float or None) – Starting date.

  • end (datetime, date, tuple, int, float or None) – Ending date.

  • active (bool or None) – Active/inactive flag. None means using its default value, which may be different for different subclasses.

  • repeater (tuple or None) – Repeater interval.

  • warning (tuple or None) – Deadline warning interval.

>>> OrgDate(datetime.date(2012, 2, 10))
OrgDate((2012, 2, 10))
>>> OrgDate((2012, 2, 10))
OrgDate((2012, 2, 10))
>>> OrgDate((2012, 2))
Traceback (most recent call last):
    ...
ValueError: Automatic conversion to the datetime object
requires at least 3 elements in the tuple.
Only 2 elements are in the given tuple '(2012, 2)'.
>>> OrgDate((2012, 2, 10, 12, 20, 30))
OrgDate((2012, 2, 10, 12, 20, 30))
>>> OrgDate((2012, 2, 10), (2012, 2, 15), active=False)
OrgDate((2012, 2, 10), (2012, 2, 15), False)

OrgDate can be created using unix timestamp:

>>> OrgDate(datetime.datetime.fromtimestamp(0)) == OrgDate(0)
True
property start: date | datetime[source]

Get date or datetime object

>>> OrgDate((2012, 2, 10)).start
datetime.date(2012, 2, 10)
>>> OrgDate((2012, 2, 10, 12, 10)).start
datetime.datetime(2012, 2, 10, 12, 10)
property end: date | datetime[source]

Get date or datetime object

>>> OrgDate((2012, 2, 10), (2012, 2, 15)).end
datetime.date(2012, 2, 15)
>>> OrgDate((2012, 2, 10, 12, 10), (2012, 2, 15, 12, 10)).end
datetime.datetime(2012, 2, 15, 12, 10)
is_active() → bool[source]

Return true if the date is active

has_end() → bool[source]

Return true if it has the end date

has_time() → bool[source]

Return true if the start date has time field

>>> OrgDate((2012, 2, 10)).has_time()
False
>>> OrgDate((2012, 2, 10, 12, 10)).has_time()
True
has_overlap(other) → bool[source]

Test if it has overlap with other OrgDate instance

If the argument is not an instance of OrgDate, it is converted to OrgDate instance by OrgDate(other) first.

>>> od = OrgDate((2012, 2, 10), (2012, 2, 15))
>>> od.has_overlap(OrgDate((2012, 2, 11)))
True
>>> od.has_overlap(OrgDate((2012, 2, 20)))
False
>>> od.has_overlap(OrgDate((2012, 2, 11), (2012, 2, 20)))
True
>>> od.has_overlap((2012, 2, 11))
True
classmethod list_from_str(string: str) → list[OrgDate][source]

Parse string and return a list of OrgDate objects

>>> OrgDate.list_from_str("... <2012-02-10 Fri> and <2012-02-12 Sun>")
[OrgDate((2012, 2, 10)), OrgDate((2012, 2, 12))]
>>> OrgDate.list_from_str("<2012-02-10 Fri>--<2012-02-12 Sun>")
[OrgDate((2012, 2, 10), (2012, 2, 12))]
>>> OrgDate.list_from_str("<2012-02-10 Fri>--[2012-02-12 Sun]")
[OrgDate((2012, 2, 10)), OrgDate((2012, 2, 12), None, False)]
>>> OrgDate.list_from_str("this is not timestamp")
[]
>>> OrgDate.list_from_str("<2012-02-11 Sat 10:11--11:20>")
[OrgDate((2012, 2, 11, 10, 11, 0), (2012, 2, 11, 11, 20, 0))]
classmethod from_str(string: str) → OrgDate[source]

Parse string and return an OrgDate objects.

>>> OrgDate.from_str('2012-02-10 Fri')
OrgDate((2012, 2, 10))
>>> OrgDate.from_str('2012-02-10 Fri 12:05')
OrgDate((2012, 2, 10, 12, 5, 0))
class orgparse.date.OrgDateScheduled(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]

Date object to represent SCHEDULED attribute.

class orgparse.date.OrgDateDeadline(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]

Date object to represent DEADLINE attribute.

class orgparse.date.OrgDateClosed(start, end=None, active: bool | None = None, repeater: tuple[str, int, str] | None = None, warning: tuple[str, int, str] | None = None)[source]

Date object to represent CLOSED attribute.

class orgparse.date.OrgDateClock(start, end=None, duration=None, active=None)[source]

Date object to represent CLOCK attributes.

>>> OrgDateClock.from_str(
...   'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] =>  0:30')
OrgDateClock((2010, 8, 8, 17, 0, 0), (2010, 8, 8, 17, 30, 0))
property duration[source]

Get duration of CLOCK.

>>> duration = OrgDateClock.from_str(
...   'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] => 0:30'
... ).duration
>>> duration.seconds
1800
>>> total_minutes(duration)
30.0
is_duration_consistent()[source]

Check duration value of CLOCK line.

>>> OrgDateClock.from_str(
...   'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] => 0:30'
... ).is_duration_consistent()
True
>>> OrgDateClock.from_str(
...   'CLOCK: [2010-08-08 Sun 17:00]--[2010-08-08 Sun 17:30] => 0:15'
... ).is_duration_consistent()
False
classmethod from_str(string: str) → OrgDateClock[source]

Get CLOCK from given string.

Return three tuple (start, stop, length) which is datetime object of start time, datetime object of stop time and length in minute.

class orgparse.date.OrgDateRepeatedTask(start, before: str, after: str, active=None)[source]

Date object to represent repeated tasks.

property before: str[source]

The state of task before marked as done.

>>> od = OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE')
>>> od.before
'TODO'
property after: str[source]

The state of task after marked as done.

>>> od = OrgDateRepeatedTask((2005, 9, 1, 16, 10, 0), 'TODO', 'DONE')
>>> od.after
'DONE'

Further resources

Indices and tables