Welcome to “The Internals Of” Online Books project! 🤙
I’m Jacek Laskowski, a Freelance Data Engineer 🧱 specializing in Apache Spark (incl. Spark SQL and Spark Structured Streaming), Delta Lake, Unity Catalog, MLflow, DSPy, Databricks with brief forays into a wider Data, ML and AI Engineering space (mostly during Warsaw Data Engineering meetups).
I’m very excited to have you here and hope you will enjoy exploring the internals of the open source projects together (in no particular order):
Please note that some books have less current content than others, but that’s expected with a one-person project where so many things are truly interesting and thus time-consuming. Life’s too short to taste everything :/
The aim of this project is to host all the current and future internals books under a single organization on GitHub and publish to a single domain via GitHub Pages (until I find a better way to publish the books).
The japila-books project uses Material for MkDocs documentation framework.
The book projects use a custom Docker image.
Use build-image.sh shell script to build the custom Docker image.
Start Colima.
colima start
Execute the build-image.sh shell script to build the Docker image.
./build-image.sh [version_tag]
Go to https://github.com/squidfunk/mkdocs-material/tags to find the available insiders tags.
docker run \
--rm \
-it \
-p 8000:8000 \
-v ${PWD}:/docs \
jaceklaskowski/mkdocs-material \
build --clean
TIP: Consult the Material for MkDocs documentation to get started.
Use docker run command with serve argument (with --dirtyreload for faster reloads) in the project root (the folder with mkdocs.yml).
docker run \
--rm \
-it \
-p 8000:8000 \
-v ${PWD}:/docs \
jaceklaskowski/mkdocs-material \
serve --dirtyreload --verbose --dev-addr 0.0.0.0:8000
Run an interactive shell in a container.
docker run \
--rm \
-it \
-p 8000:8000 \
-v ${PWD}:/docs \
--entrypoint sh \
jaceklaskowski/mkdocs-material
While inside, execute the following command to list outdated packages, and show the latest version available (as described here).
pip list --outdated
Use scripts/sync-version.py to bump the version of the project a book is about (e.g., extra.spark.version in mkdocs.yml) and sync the dependency versions (extra.*) from a local git checkout of that project.
Execute the script in the book project root (the folder with mkdocs.yml).
../japila-books.github.io/scripts/sync-version.py v4.2.0 ~/oss/spark
Use --dry-run to print the changes without writing mkdocs.yml, and --latest (in place of the version) to use the latest release of the project.
../japila-books.github.io/scripts/sync-version.py --latest ~/oss/spark --dry-run
A book can fine-tune the sync with an optional scripts/sync-version.json config (e.g., to pick the highest of several matching versions or to build URLs from a version). Use --help for the details.
Use scripts/tag-release.sh to create a git tag matching the version of the project a book is about (e.g., extra.spark.version in mkdocs.yml) and push it to origin.
Execute the script in the book project (mkdocs.yml is looked up in the current directory and its parents).
../japila-books.github.io/scripts/tag-release.sh
The project is auto-detected as the extra.<key> block in mkdocs.yml with a version and a github: https://github.com/<org>/<repo>/blob/<ref> field (e.g., extra.spark, extra.delta, extra.uc). Use --key to pick one explicitly and --yes to skip the confirmation prompt.
../japila-books.github.io/scripts/tag-release.sh --key spark --yes
The script requires yq.