How do I make sure that my awesome, glorious program is correct? Good question, and the short answer is: in real life you don’t. Especially if your code is (even moderately) complex—not even mentioning whether it is asynchronous or is meant to communicate with hardware.
For compiled high-level languages, at the very least, the compiler makes it for a first barrier for all obvious (and a good deal of less-than-obvious) mistakes—if you get a type wrong, just to name one, it will happily refuse to move forward. Unfortunately Python is interpreted and this kind of safeguard is just not there: the interpreter will exit upon an syntax error, but that’s pretty much it. All the rest is caught at runtime. If you are in production, that simply means too late.
Just so that we do not focus too much on the negative side of things, this is not to say that there is nothing you can do. In fact, besides paying attention, there are at least two things that you can do, even in interpreted languages:
unit testing; and
static analysis.
Generally speaking, physicists hate both, and profoundly too. But if you are writing code, they should come right next to version control in your work-flow toolbox. This is exactly what this chapter is about.
9.1 Unit testing
Unit testing is about testing that your code does the right thing in a well defined set of specific conditions where you know the answer in advance. Let us illustrate this in a trivial case:
def square(x):"""Toy function returning the square of x. """return x**2.def test():"""Dumb test battery. """assert square(2.) ==4.assert square(2) ==4.assert square(-2.) ==4.print("All tests passed!")if__name__=="__main__": test()
All tests passed!
This makes it for a sensible structure, at least in principle: here is my small Python module, exporting the function square(), and a small test harness that I can execute by simply invoking the file from the terminal. I have a few, representative cases covered, here: positive and negative floats, as well as an integer. (It goes without saying, we can always add more.) If everything goes well, I get a nice “All tests passed!” in response.
Note
In Python, the keyword assert checks that a condition is true, and raise an AssertionError exception otherwise. (You will learn what is an exception in Chapter 8.) Note assertions can be disabled by running Python with the -O (for optimized) command-line switch, so you should be careful what you use the assert for.
9.1.1 Unit testing basics
We are on the right track, but not quite there, yet. In order to make unit testing really effective we need a sensible set of design rules and good practices, as well as some logistical infrastructure to cover the practical side of things.
First and foremost, break up your program in small, isolated pieces (e.g., functions or methods), each encapsulating a well-defined and (if at all possible) simple functionality, and test each piece separately. In practice, this is accomplished by means of a sensible hierarchy of functions and classes—which is often the hardest thing to get right when you structure your code. If you make sure that each single piece behaves as expected in the largest possible set of external conditions (e.g., with different input arguments) then you add confidence that the whole thing behaves as intended too. This is generally much simpler than testing the whole program at once—which you will have to do, too.
No matter what you do, if your code is actively used, you can be sure that it will change with time—in response to bugs that users will find, or new features they will request, or even in response to the natural evolution of the underlying language. If your codebase is (even moderately) complex, each change, no matter how small it is, will have a huge phase space for breaking things in unrelated places, through ramifications that you did not anticipate. If you have an extensive suite of unit tests, chances that you catch any possible problem early increase. You cannot be 100% sure, but, generically speaking, your confidence increases with the coverage of your unit tests.
What does make a unit test a good unit test? Well, it is hard to tell on purely speculative grounds, but there are some common features that you typically would like to have:
test a very specific behavior with a binary output—a test should pass or fail;
run as quickly as possible, since in the long run you will want many of them;
be deterministic, i.e., produce the same result every time, even when you need a PRNG;
cover normal cases, edge cases, and errors; and as many as possible.
Reality is, physicists tend to work with a development model that one might call physicist-driven development (PDD), i.e.:
write a whole bunch of code that does something—it doesn’t matter how;
swear you will come back to it tomorrow and write the unit tests;
add more code, doing yet more things, and more complicated.
Professional programmers, on the other hand, are more prone to what generally goes under the name of test-driven development (TDD), which sounds more like:
write an empty placeholder for your new function;
write all the unit tests that you can imagine—they will all fail, initially;
implement your function and tweak it until all the tests pass.
I’ll let you pick which one makes more sense. TDD is not mandatory, and there are alternative, viable ways to do business. PDD is a recipe for disasters. Period.
9.1.2 Unit testing in practice
Back to the logistical side of things. The snippet of code we have shown at the beginning of this section, if admirable in spirit, comes with a number of shortcomings. Most notably, everything happens manually: you have to run the script yourself, and you have to inspect the output yourself. As your code grows in complexity, this is just not sustainable. Unit testing should happen automatically, all the time, and somebody or something should come get you out of bed whenever a test fails.
The Python standard library provides a module called unittest. Just like the logging module, it is one of the few modules in the standard lib that did not withstand the test of time and, while kept for backward compatibility, it is rarely used for new projects.
If you are starting fresh, you might as well opt for the third-party package pytest, as everybody else is doing these days. From a practical standpoint, this means creating a folder in your repository called tests and move all your test code in a series of dedicated Python modules in there, which each module being a series of loose function whose name starts with test_
from mymodule import squaredef test_square():"""Unit test for the square function. """assert square(2.) ==4.assert square(2) ==4.assert square(-2.) ==4.
You might have notices that we have removed the final print() statement: there is no need for it, since it is now pytest that is responsible for collecting and running the tests, and examining the output. How? Simple:
pytest
I know: it sounds magic, but it is really that simple. All you have to do is concentrate on writing the tests, and pytest will take care of all the bookkeeping.
Ah, one last thing: you are welcome to trigger the tests manually, but, as we said, unit tests should run automatically no matter what you do. This is sometimes referred to as continuous integration, and the standard way to achieve that in 2026 is through GitHub or Codeberg actions. The basic idea is that under specific conditions (e.g., when you push a change on the main branch, or on a branch with a pull request) a complete run of the unit tests is triggered and you get an email if something goes wrong.
If you want to really see how all these things play out in practice, go ahead and look (again) at metarep—make sure to hit the links to the unit tests and the GitHub actions.
How do you integrate all of this in your daily workflow? Easy! Did you just find a bug in your code? Make sure you add a unit test along with the fix, so that you will never be hurt again by that particular bug. (There will be others, of course.) Are you adding a new feature? Make sure the new code is covered by unit tests.
9.2 Static code analysis
By its very nature, as we said, Python will show you all the errors at runtime. Now, say you have a bug in a part of the code that is exercised very rarely, and not covered by unit tests: the very first time that you do exercise that particular path, the Python interpreter might crash, or happily do something that is not what you intended it to do. It might take years for even realizing that there was a bug in the first place.
If you think about, many common mistakes can be found by just looking at the code and without running it. And in fact you might argue that all of them can, at least in principle, provided that your sight is sharp enough. The point is: a good part of it can be done programmatically. Generally, a program will not understand your program, but a program can be designed to spot specific kinds of errors and inconsistencies. This is customarily called static code analysis; pylint and ruff are popular tools doing just that.
Note
For good or bad, this is becoming more and more uncertain as the AI revolution unfolds and LLMs become better and better at understanding code and finding mistakes. While this is true in many different areas, the implications vary. Unit tests will presumably continue to be relevant for the foreseeable future: you might have a horde of agents writing unit tests for you, but you will always need unit tests—which means that you will presumably have to understand what a unit test is. On the other hand, deterministic code analysis might fall out of fashion in favor of AI-driven code review, and this section of the notes might become obsolete. We’ll see.
Take the following snippet, for example.
x =1.y =2.very_uncommon_condition =Falseif very_uncommon_condition:print(x + z)else:print(x + y)
3.0
See anything wrong? If we entered the first branch of the if statement we would have crashed because we mis-typed y as z.
Let’s feed this into pylint, and see what happens
[lbaldini@nbbaldini ~]$ pylint snippet.py
************* Module snippets.linting1
snippet.py:1:0: C0111: Missing module docstring (missing-docstring)
snippet.py:1:0: C0103: Constant name "x" doesn't conform to UPPER_CASE naming style (invalid-name)
snippet.py:2:0: C0103: Constant name "y" doesn't conform to UPPER_CASE naming style (invalid-name)
snippet.py:3:0: C0103: Constant name "very_uncommon_condition" doesn’t conform to UPPER_CASE naming style (invalid-name)
snippet.py:5:14: E0602: Undefined variable 'z' (undefined-variable)
--------------------------------------------------------------------
Your code has been rated at -5.00/10 (previous run: -5.00/10, +0.00)
For one thing I did not follow PEP8 and got some unfortunate variable name. This is excusable, but the worst (or best, depending on how you look at it) part is that pylint did find the typo! It also assigned a score (\(-5\) out of \(10\), for the record) to the overall coding.
We are not going to drag this any longer, but you should consider using static code analysis routinely as you program in Python. Static analysis tools tend to be quite verbose, it is true, to the point where they can become annoying; in addition, they try and enforce many different (good) things at once:
formal correctness;
efficiency;
avoiding anti-patterns;
style guides.
But they are also typically highly customizable, and you can can mute errors you don’t care about with only a few lines in your pyproject.toml file. (See here for an example.) Including some static analysis in your continuous integration, for instance, is often a good idea. Up to you to strike the right balance.