Category: Level 1 Undergraduate

  • Complex numbers: an introduction

    Complex numbers: an introduction


    Complex numbers have fascinated me since high school. Usually, it’s where we are taught about natural numbers, integers, rational, irrational, and real numbers but never about complex numbers. This post is for those who might be interested in an easy introduction into the realm, or rather, plane of complex numbers. And they’re not without practical significance either: no electronic device such as the one you’re using to read this post could have been built without physicists, electrical engineers, and computer scientists knowing anything about the gift of complex numbers from sixteenth century mathematicians.

    Blown away

    ‘There are such things as negative numbers’, explained my father to me when I must have been about six or seven years old since I was a second-year pupil in primary school. He explained the notion of negative possession when owing a certain number of marbles to someone which was greater than the number of marbles you physically carry with you. As this was one of those I-still-remember-where-I-was-when moments, like it was yesterday, I remember sitting on the floor besides the coffee table in the living room of our terraced house in the town of Emmeloord, which had been reclaimed just forty-three years earlier from the IJsselmeer, a lake formerly part of the North Sea.

    I clearly remember feeling exactly the same when he had told me earlier our planet wasn’t flat and when my Mum told me in the car yet a few months earlier, that we were living on the sea floor. The cap of my mind was blown away, yet again. It took a while before I managed to fold my slow and wet brain lobes around the notion that negative numbers existed, even though you couldn’t see them in the real world like you could ‘see’ regular numbers such as in lengths or the number of marbles(beginfootnote)Inexplicably, I had never considered the fact that temperature could get below 0 ℃, which it still did, back then in The Netherlands. We used to enjoy an outdoor activity called ice skating, on frozen lakes, ponds, rivers, and ditches.(endfootnote).

    I hastened to tell my primary school teacher excitedly about negative numbers. She just nodded and then told me to proceed with doing my homework on boring regular arithmetic. She had a point as I wasn’t very good at it.

    Fast forward to when I must have been about fifteen or sixteen when I read about complex numbers in a popular textbook about quantum mechanics. The fact that they were called ‘complex’ may have triggered my curiosity as I assumed that term pertained to it being very difficult, but mostly because, apparently, so-called imaginary numbers are a thing! I had that exact same feeling again. The cap of my mind had melted. The whole notion seemed to radiate some kind of magical power. What sorcery was this? Could this be a doorway to extra dimensions?

    The next day, I told my mathematics teacher, Mr Es – Es is not his actual name but it was his two-letter code in our high school timetable. I’ve always found it appropriate Es is also the symbol for the element Einsteinium in the periodic system. As his first name happened to be the same, my friends and I used to joke that we were on our way to the lessons of Albert Einstein.

    Mr Es did what every good teacher does when a student tells you something they get enthusiastic about: he encouraged it – in his case by lending me his old textbook from when he was a first-year mathematics student in Amsterdam. It was an introductory text about complex numbers at the level of undergraduate mathematics.

    The very textbook. (Click to enlarge.)

    I’m ashamed to say I kept it. It was one of those instances where, after the nth time of moving house, I realised, oh my god, I still have this!? It’s also true that I treasured it. It carries a special meaning to me. It signifies how, at least once in my lifetime, I felt acknowledged in what stirred me deeply at the time. A thing I couldn’t really share with friends or anyone close in general, I suddenly shared with someone very clever whose name was denoted by the symbol for Einsteinium.

    Thanks to the miracle of internet, we got back in touch, about twenty-five years later. I confessed I had always kept it and apologised. He had indeed wondered where it had been as he once wanted to show it to someone else. But I could keep it as he was cleaning out the attic anyway. And he was glad it had done something for me as he learnt about my current engagements in a bit of maths and physics.

    I felt guilty. I still do. Someone else could have enjoyed it just as much as I have. And now I have prevented that from happening through his book. So, whoever you are, my sincerest apologies.

    I hope, one day, I will be able to ignite sparks of joy for the beautiful mathematics of complex analysis to many others. I also hope you might experience at least a fraction of the amazement I felt and that the newly gained insight on the concept of ‘numbers’ might turn out to be beyond what you were able to imagine so far. So, let this be a beginning.

    Number sets

    A game of hopscotch drawn on the pavement with numbers on the tiles

    We all know and love (or hate, depending) the natural numbers: the whole numbers we count things with. 1, 2, 3, etc. Some mathematicians will want to include the number 0 while others don’t. In any case, this mathematical set of numbers is called the natural numbers and is denoted by the symbol $\mathbb{N}$.

    Then my father told me about the negative numbers, such as -1, -2, -3, etc. If you include the natural numbers and add to that these negative numbers, and add the number 0 to it (if you hadn’t already), then the result is an entirely new set of numbers called the integers, denoted by the symbol $\mathbb{Z}$.

    To denote that the set $\mathbb{N}$ is part of the larger set $\mathbb{Z}$, people use this symbol for subset, $\subset$. They will write $\mathbb{N}\subset\mathbb{Z}$, the natural numbers are a subset of the integers.

    Of course, there’s the ratio’s. The fractions. Between 1 and 2, there’s 1.5. So, in fraction-notation, that’s $\frac{3}{2}$. They’re obviously not whole numbers. They’re rational numbers because they can be represented by a ratio of integers. This number set is symbolised by $\mathbb{Q}$. We now have $$\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}.$$

    It is interesting to note that, therefore, by this expression of subsets of subsets, even numbers such as 9 are rational numbers. On the surface, it’s not a fraction. Below the surface, however, it can be expressed as a ratio of integers: $9=\frac{9}{1}=\frac{18}{2}=\frac{36}{4}$, for example (and infinitely more).

    But wait, there’s more. Fractions such as 1.5 and 3.2 are finite. What if the decimals don’t end? What if you can’t write a particular kind of numbers as ratios, such as with the number $\pi$ or $\sqrt{2}$? These numbers are called the irrational numbers. They are all the numbers which aren’t rational. There’s no symbol for that(beginfootnote)Often, mathematicians circumvent the lack of a symbol by writing something like ​​​$\mathbb{R} \backslash \mathbb{Q}.$(endfootnote).

    Instead, there’s a symbol for all the natural numbers, the integers, the rational numbers, and the irrational numbers altogether(beginfootnote)Yes, indeed, my dear fellow mathematician, you thought correctly, I am skipping transcendental numbers here (and algebraic numbers, for that matter). As all transcendental numbers are irrational numbers but not all irrational numbers are transcendental, I decided it over-complicated things in what was supposed to be an introductory text on complex enough numbers anyway.(endfootnote). They’re called the real numbers and this set is denoted by $\mathbb{R}$. This is the set we’re all used to working with. We now have $$\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}.$$

    The set of real numbers $\mathbb{R}$ contains all the numbers. Or does it?

    A diagram of all the number sets in the shape of ellipses. The ellipse of R containing the ellipse of Q containing the ellipse of Z containing the ellipse of N.

    The secret of del Ferro, del Fiore, Tartaglia, and Cardano

    Well, you guessed it. Here they come, the complex numbers. Let’s do just a tiny bit of maths. Remember what the quadratic of a number was? And what a square root was? What is the square root of 64, in other words, $\sqrt{64}$? Yes, that’s 8. Because 8 times 8, or 8 squared, or $8^2$ equals 64.

    Okay, suppose $x^2 = 64$, what is $x$ then? Well, you do exactly the same thing, you un-square $x$ by taking its square root. And you have to do the same with the number after the equal sign. So, $\sqrt{x^2} = \sqrt{64}$, in other words, $x = 8$.

    Tartaglia

    Maybe you remember this comes in handy when calculating the lengths of the edges of your piece of land. Suppose, the surface area of your square piece of land is 64 square kilometre (or square miles). What is the length of an edge of that land? That’s 8 kilometre (or miles).

    All these calculations take place in the realm of $\mathbb{R}^+$, the positive part of all real numbers. Note that no surface area of a piece of land can be negative. In other words, a surface area of -64 square metres is nonsensical. Also, the square root of -64 has no solution. It’s not -8, because -8 times -8, or $(-8)^2$ is simply 64 again, because a negative number times a negative numbers equals a positive number as we proved in an earlier post.

    Cardano

    Sometime in the sixteenth century, somewhere in Italy, Scipione del Ferro, professor of the University of Bologna, solved a slightly different kind of equation. It was a so-called cubic equation. Where we basically found the solution to a quadratic equation such as $x^2 = 64$ from the top of our heads, he found solutions for a cubic equation such as $x^3 + x^2 + 6x + 3 = 0.$ Del Ferro was known for not wanting to publish any of his proofs and solutions. He kept a secret notebook and that was it.

    On his death bed, however, he told his pupil Antonio Maria del Fiore the secret to solving it. Del Fiore went on to challenge Niccolò Fontana Tartaglia, a mathematician residing in Venice at the time. Tartaglia had actually solved it himself before and trusted the formula to Gerolamo Cardano, the then Milan-based polymath and genius. Tartaglia messaged the solution in the form of a poem (no less!) but didn’t entrust the proof to him.

    Of course, Cardano was able to reconstruct the proof anyway. As he learnt that del Ferro had also found the solution, he then proceeded to publish it all in his Ars Magna from 1545, much to the chagrin of Tartaglia.

    So, what was the secret so many large minds had been secretive about? A new type of number.

    imaginary

    Let’s take a simpler example. Suppose, we have the following simplistic quadratic equation: $x^2 – 4 = 0$. To solve it, we ‘move’ the 4 to the other side of the equal sign, by adding 4 to both sides: $x^2 – 4 + 4 = 0 + 4$, which simply becomes $x^2 = 4$. If you apply the square root to both sides, you get $\sqrt{x^2} = \sqrt{4}$. The solution to this equation is thus $x=2$ or $x=-2$ (because $-2\times -2 = 4$ too).

    Good. Basically, the mathematicians of the sixteenth century opined that they should be able to solve a variation of this equation as well: $x^2 + 4 = 0$. Let’s bring the 4 again to the other side of the equal sign by subtracting 4 on both sides: $x^2 + 4 – 4 = 0 – 4$, which becomes $x^2 = -4$. Now, again, the question is, what is $x$?

    Let’s try and apply the square root to both sides again: $\sqrt{x^2} = \sqrt{-4}$. Halt. Stop. What is the square root of -4? What is the square root of a negative number?

    We have the same situation where we are to apply the square root of a negative surface area. The answer isn’t -2, because $-2\times -2 = 4$, not -4. What then?

    Before del Ferro, Tartaglia, and Cardano, people would have said that there simply is no solution. Thanks to them, however, we can solve it. The answer lies in the following definition: $$i^2=-1.$$

    This seemingly simple act enables us to solve $x^2=-4$. We can then write $x = 2i$ or $x = -2i$.

    Let’s take our first solution, $x = 2i$. If we square this, we get $x^2 = (2i)^2$, which we can also write as $x^2 = 2^2i^2$. Now, since $i^2 = -1$, we can substitute that to get $x^2 = 2^2(-1)$, which is, of course, $x^2 = -4$. Ecco!

    The same goes for the other solution, $x = -2i$. If we square this, we get $x^2 = (-2i)^2$, which we can write as $x^2 = (-2)^2i^2 = 4i^2 = 4(-1) = -4$. Ecco!

    So, you may ask, what devilish entity is this $i^2=-1$? The letter $i$ stands for ‘imaginary’ and so, $i$ is a so-called imaginary number.

    Now, because $i^2=-1$, you can also write(beginfootnote)Although, I actually prefer to use $i^2=-1$ over $i=\sqrt{-1}$ even though the latter has been mentioned in many school books. However, I believe it might lead to confusion. Since we have the rule that $\sqrt{a}\sqrt{b}=\sqrt{ab}$ where $a$ and $b$ are positive real numbers, you might try to apply this rule to negative real numbers, such as when $a=b=-1$. You would then get the incorrect statement $\sqrt{-1}\sqrt{-1} = \sqrt{(-1)(-1)} = \sqrt{1} = 1$, which is wrong as it should be equal to -1. That’s why I try to avoid using $i = \sqrt{-1}$ where I can.(endfootnote) that $i = \sqrt{-1}$. And that’s the crazy part: how can you calculate the square root of a negative number? How can you calculate the square root of a negative surface area? The answer is, you can’t. Not in the realm of the real numbers $\mathbb{R}$, that is. However, we’re not in Kansas anymore, Dorothy. We’re in a new land called the complex numbers. Bye $\mathbb{R}$, and welcome to $\mathbb{C}$.

    Here are some examples of complex numbers: $2i$, $\frac{2}{3}i$, $i\sqrt{2}$, $i \pi$, $-0.25i$. What’s more, you can add these imaginary numbers to a real number such as 3, like so: $3 + 2i$ or $3 + \frac{2}{3}i$ etc. These sums are their own answer. They are complex numbers.

    A complex number $z$ is of the form $z = a + bi$, where $a$ and $b$ are real numbers and $i^2 = -1$. The first real number, $a$, is called the real part of $z$. The last real number, $b$, is called the imaginary part of $z$. The set of all complex numbers is denoted by $\mathbb{C}$.

    And so, we now have

    $$\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}\subset\mathbb{C}.$$

    Note that every real number is a complex number but not every complex number is a real number. That is what one thing being a subset of another thing means. For instance, the real number 9 is a complex number where $b=0$. In other words, the real number 9 can be written as the complex number $9 + 0i$, which is simply 9, which thus happens to be a real number too.

    But $z = 3 + 2i$ is not a real number, because it has an imaginary part which is not equal to zero. So, $z$ is now exclusively a complex number.

    A diagram of all the number sets in the shape of ellipses. The ellipse of C containing the ellipse of R containing the ellipse of Q containing the ellipse of Z containing the ellipse of N.

    Complex plane

    Graphically, all the real numbers of $\mathbb{R}$ can be thought of as a point on the number line.

    A diagram depicting the real number line. Every point on this line represents a real number, such 0, 1, 2, 3 and the square root of 2, pi, and e.

    So, where do complex numbers reside?

    Owing to people such as Wallis, Wessel, Argand, Buée, Mourey, Warren, Français, Bellavitis, Gauss, and Euler[1], the idea to extend the real number line with an imaginary number line perpendicular to the real number line came to fruition. What you get is the so-called complex (geometric) plane, sometimes called the $z$-plane, Gauss plane or Argand plane.

    So, a complex number such as $z = 3 + 2i$, ‘contains’ the real number $3$ along the real axis, and the imaginary part, along the imaginary axis, sits at $2i$. A complex number is therefore always represented by a point in a two-dimensional space. Note that all the numbers from all the subset of complex numbers, i.e. $\mathbb{R}$ all the way down to $\mathbb{N}$, can also be represented by a point in this same two-dimensional complex space – it’s just that they all reside on the real axis.

    As you can -heh- imagine, doing calculations with complex numbers has become an exercise of geometry now! In fact, one of the most beautiful equations in mathematics (at least to my taste) pertains to trigonometry in the complex plane; it’s called Euler’s Formula.

    A diagram representing the complex plane. Perpendicular to the real number line is now a so-called imaginary axis with numbers such as i, 2i, 3i, pi-i, i square root of 2, etc. A complex number is now a point in on that surface.

    Not so imaginary

    It’s unfortunate that this number $i$ and any real number multiplication of it are called imaginary numbers. It was the renowned French philosopher and mathematician René Descartes who coined the term imaginary numbers because he considered them to be illusory. In fact, even Cardano had described them as ‘some recondite third kind of thing’[2].

    It’s unfortunate because ‘imaginary’ leads to semantic ambiguity. I get it: you would never see something like $\sqrt{-1}$ in the real world. But neither would you see $\sqrt{2}$ out in the wild, for that matter. And yet, it’s the exact length of the hypotenuse of a particular right triangle, which a skilled DIY person could make while you’re waiting. To me, ‘real’ numbers such as $\pi = 3.1415926535897 \dots$ without ever ending are as real as ‘imaginary’ numbers are (and vice versa). Circles are a real thing and $\pi$ can be used to do calculations on them. Well, with imaginary numbers you can do calculations on them just as well.

    Complex numbers are used in a variety of sciences. In Einstein’s relativity, which makes GPS navigation possible, you could make use of so-called imaginary time. This sounds like a concept straight from a science-fiction novel, however, imaginary time is a well-defined concept. In fact, in a previous post, we used this to derive the central set of equations in relativity, called the Lorentz transformations. See how the word ‘imaginary’ might invoke unwanted ambiguity?

    To make quantum mechanics work – the most successful theory to date – complex numbers are all over the place. Without them, the computer, mobile phone, tablet, TV, VCR, even your modern fridge – they wouldn’t have worked as no engineer would have been able to produce integrated circuits. The wave function is a complex function living in a complex separable Hilbert space, taking on complex probability amplitudes, evolving according to the Schrödinger equation, which itself is a complex equation.

    In mathematics, one of the better-known areas of research where complex numbers play a central role is the study of complex dynamical systems. The featured image above is a detail of the famous Mandelbrot set. It’s a special collection of complex numbers, the projection of which you see plotted colourfully in the complex plane. The study of (complex) fractals also informs all kinds of patterns in nature and growth, even weather forecasts, and climate science – they’re all informed by complex-dynamical areas of mathematical interest. Also, we’ve used them in a previous post, calculating whether a lab centrifuge with $n$ available spots can be balanced out by a $k$ number of test tubes.

    A fun application of complex numbers is computer games. To calculate rotations in three-dimensional space, computer scientists make use of quaternions, which are an extension of the complex plane. A quaternion is an expression of the form $a + bi + cj + dk$, where $a,b,c,d$ are any old real numbers, and $i^2=j^2=k^2=-1$. However, this is perhaps an interesting subject for another bit of maths and physics.

    [1] Cooke, R. (2005) The history of mathematics : a brief course. 2nd edn. New York, N.Y.: Wiley.

    [2] Open University (2014) Essential mathematics 1. Milton Keynes: Open University.

    Images

    Featured image: Mandelbrot set – Step 6 of a zoom sequence by Wolfgang Beyer under CC BY-NC-SA 2.0; adapted to fit layout.

    Hopscotch Game by ncassullo.

    Niccolò Fontana Tartaglia. Rijksmuseum, Dutch National Museum. Public domain.

    Girolamo Cardano. Wellcome Images under CC BY 4.0.

  • Heisenberg’s uncertainty principle

    Heisenberg’s uncertainty principle


    It’s perhaps not as famous as Einstein’s formula but in this day and age many people may still have heard at least once of the phrase ‘Heisenberg’s uncertainty principle’. It plays an important role in quantum mechanics. You may have heard that every time you observe or measure matter, due to the crudeness or inherent inaccuracy of the measurement device, you will inevitably disturb your own observation. This would then preclude you from gaining accurate knowledge with satisfying certainty. In fact, in general, Heisenberg’s uncertainty principle states that nothing can be certain. At the risk of sounding vague and vanilla, all of these statements are completely and utterly wrong. Let’s look at what it really says, shall we?

    Figure 1. Werner Heisenberg in Göttingen in 1924.

    Fourier transform pairs

    Trade-offs. Who doesn’t hate them? Remember when your parents told you that you could have this but then not have that or maybe just a bit of this but then less or fewer of that? Unsurprisingly, at least three famous philosophers have written a few words on this, each in their own way lamenting on the existence of trade-offs and how to deal with them. One chose to become all rebellious about it and wrote: ‘I want it all, I want it all, and I want it now!’ (May, 1988). The other two, however, chose to be more pragmatic about it as they postulated that ‘you can’t always get what you want’ (Jagger & Richards, 1968). Obviously, they knew that, sometimes, life brings you Fourier transform pairs. The more well-known example is of course Heisenberg’s uncertainty principle.

    If you limit a particle’s range of possible positions in space $(\Delta x)$, you increase its range of possible momenta(beginfootnote)Momentum is the product of mass $m$ and velocity $v,$ so $p=mv.$ It’s a measure for the amount of motion of an object.(endfootnote) along the $x$-direction $(\Delta p_x),$ and vice versa.

    This is formalised as follows:

    $$\Delta x \Delta p_x \geq \frac{\hbar}{2}.$$

    Just to be absolutely clear: the delta-symbol $\Delta$ is a range of a certain quantity. Usually, a $\Delta$ is defined as the difference between two values. Suppose, you measure point $A$ of your garden fence to be $0.1$ metre away from your wall and point $B$ to be $0.7$ metre away from your wall, then the $\Delta$ of the distances, i.e. the length between points $A$ and $B,$ is $0.7-0.1=0.6$ metre.

    In Heisenberg’s principle, it is stated that the product of the range of possible positions $\Delta x$ and the range of possible momenta $\Delta p_x$ is greater than or equal to some number. Mind you, it’s a tiny number. The symbol $\hbar$ stands for the Planck constant divided by $2 \pi,$ and the result gets cut in half yet again.

    This means that whenever one is getting bigger, $\Delta p_x$ for instance, the other is getting smaller, which is then $\Delta x.$ And vice versa.

    Click here if you’d like to do a bit of maths. It’s very easy.

    Just to get an intuitive insight in this relation, suppose $\frac{\hbar}{2}=1,$ and so, suppose, $\Delta x \Delta p_x = 1.$ Furthermore, suppose $\Delta x = 0.5.$ What value does $\Delta p_x$ has to be to satisfy this equation? Exactly, $\Delta p_x$ has to be $2,$ because $0.5 \times 2 = 1,$ or else the equation is false.

    Now, lets make $\Delta x$ smaller. In other words, we’re going to try to pinpoint the location with much more precision. So, let’s say, $\Delta x = 0.001.$ What value does $\Delta p_x$ has to become to satisfy this equation? You guessed right, $\Delta p_x$ has to become even larger: $\Delta p_x = 1000,$ because $0.001 \times 1000 = 1.$ If you were to reverse the situation – decreasing the size of $\Delta p_x$ – then, in turn, $\Delta x$ would have to become larger.

    In reality, $\frac{\hbar}{2}$ is much smaller than 1. It is, in fact, about $5.273 \times 10^{-35} \text{J/s}.$ That’s thirty-four zeros behind the decimal point and then ending in 5273. It’s incredibly small. Don’t worry about this. We’ll get back to that later.

    Hopefully, now you see the relation between $\Delta x$ and $\Delta p_x$ as put forward by Heisenberg’s formulation. They complement each other. Whenever a range of possible values becomes larger, in other words, the $\Delta$ or range of value-options is larger – its actual value becomes more uncertain, hence the use of the word ‘uncertainty’ in Heisenberg’s uncertainty principle(beginfootnote)In fact, it’s statistics. The $\Delta$-sign could just as well be a $\delta$-sign, so $\delta x \delta p_x \geq \frac{\hbar}{2},$ which signifies its statistical character more accurately. After all, the wave function is about probabilities.(endfootnote).


    But why is this? While this principle plays a central role in quantum mechanics, it’s actually not fundamentally a quantum-mechanical law. This principle exists more generally in many instances in physics, and, even more generally, in mathematics.

    In mathematics, the variables position and momentum are said to be a Fourier transform pair. Put in yet other mathematical jargon, position and momentum are said to be conjugate variables.

    Sound

    A well-known, non-quantum-mechanical example of the uncertainty principle is determining the pitch of a sound. How ‘high’ a note is, depends on the frequency.

    The most familiar way we depict sound waves is a simple sine wave. It represents the simplest of sounds possible. Also, it’s the most boring of sounds possible.

    The $x$-axis represents time. The $y$-axis represents the amplitude of the sound or the loudness, the intensity of it. As you can see, the sound wave repeats itself over time; the pattern is cyclic. One whole cycle is when the plot has completed going up, going down, going further down, and going up again. The time it takes to complete one cycle is designated by the symbol $T,$ called the period(beginfootnote)It is also possible to measure the time-distance between two peaks or two troughs.(endfootnote). So, this particular sound wave is said to be periodic.

    Figure 3. A time-amplitude plot of a boring old sinusoidal sound wave. (Click to enlarge.)

    The shorter the period ­– the quicker the cycles are – the higher the tone. Another way of saying, is that the higher the frequency, the higher the tone. The mathematical relationship between period $T$ and frequency $f$ is the following expression:

    $$f = \frac{1}{T}.$$

    If the period gets shorter, i.e. the value of $T$ becomes smaller, then the value of $f$ becomes larger, which means higher, which means a higher tone.

    Seeing as the time period $T = 2 \pi$ seconds, the frequency diagram looks like a spike at $\frac{1}{2 \pi}$ Hz. In this frequency-diagram, the $x$-axis is the frequency and the $y$-axis is still the amplitude.

    Figure 4. A frequency-amplitude plot of the sound wave of Figure 3. It shows the exact frequency at which that sound wave exists.

    So, there are now two ways in which we can describe the sound wave: either by frequency (Figure 4) or by change over time (Figure 3).

    Notice that the sound wave plotted as a function of time (Figure 3) has no beginning nor end. For all we know, that plot could just go on forever, to an infinite amount of time, in both directions. Suppose, we would ask the question at what time exactly does the sound exist? The answer is: always. There is no particular, specific time at which it exists.

    In other words, we could write that $\Delta t = \infty.$

    Notice, however, that the frequency plot looks very finite: just one stroke. One well-defined, finite stroke. If we were to ask the question what frequency exactly does the sound have? The answer is: there is a particular, specific, exact frequency at which it exists and it is $\frac{1}{2 \pi}$ Hz ​
    $( \approx 0.16).$

    Fourier analysis

    In reality, no sound is going to be infinitely long. Pluck a guitar string and it will fade out as the energy dissipates slowly. Also, at some point it started ­– meaning, before that, it didn’t exist. In other words, in reality, a sound wave usually exists in a finite range of time.

    Let’s limit our sound wave to a range in time, so it looks more like the sound of a ‘blip’ and less like an infinite tone of boredom. Again, the $x$-axis represents time and the $y$-axis represents the amplitude.

    Figure 5. A time-amplitude plot of a so-called wavelet, a short sound burst. Contrary to the sound wave in Figure 3, it’s not infinitely long. It’s now also more difficult to assess its frequency.

    As you can see, the sound now exists in a more defined range of time – roughly 1.5 seconds. In other words, $\Delta t \approx 1.5$ seconds. That’s a whole lot smaller than the old $\Delta t = \infty.$

    Now, we ask ourselves, what is its frequency? The difficulty now is that it’s hard to pinpoint an exact period $T$. The evolution of the plot is quite different from our infinitely long sine wave. Yes, we can identify kind of those cycles we’re looking for, however, no cycle has the same shape, so, technically, we’re dealing with multiple cycles at once. And guess what, its frequency-amplitude plot looks like this.

    Figure 6. The frequency-amplitude plot of the wavelet in Figure 5. It’s far from being a specific, exact frequency. At varying degrees, it’s actually a few frequencies at the same time.

    As you can see, it has become difficult to pinpoint the exact frequency of our wavelet. It exists at a variety of frequencies and amplitudes.

    So, while the ‘time window’ of the sound wave has become more exact, the frequency has now become ‘less certain’.

    The brilliant mathematician Joseph Fourier discovered that a wavelet such as in Figure 5 can actually be constructed by adding many infinite waves at many frequencies. Put differently, Fourier analysis shows that our wavelet is the culmination of a superposition of many waves at many frequencies.

    Figure 7. The wavelet at the bottom is constructed by many infinite waves at many different frequencies superposed onto each other. This automatically means that the wavelet’s exact frequency is fundamentally harder to determine than the frequency of the sound wave in Figure 3.

    Now you see why the frequency-amplitude plot has changed from a very specific value in Figure 4 to the wider set of frequencies in Figure 6. In the latter case, the wavelet ‘contains’ multiple waves at multiple frequencies, so when you Fourier transform its time-amplitude plot to its frequency-amplitude plot, the frequency has become ‘uncertain’.

    The relation between time $\Delta t$ and frequency $\Delta f$ in ordinary classical physics is fundamentally complementary. No quantum mechanics needed.

    In mathematical jargon, time and frequency are so-called Fourier transform pairs or conjugate variables.

    The term ‘Uncertainty principle’ pertains to the general phenomenon that Fourier transforms (such as between time and frequency) entail a fundamental, mathematical trade-off between types of information carried by the two transformed variables. Heisenberg then showed that this principle also holds in quantum mechanics. And so, the uncertainty principle in quantum mechanics is called Heisenberg’s uncertainty principle.

    The De Broglie relation

    Time to go back to quantum mechanics. Remember that a particle’s best description is a wave function? A wave function is the mathematical expression of a particle containing all possible states it can assume once we measure it.

    Instead of a time-amplitude plot, let’s represent a particle by a space-amplitude plot. To make it a little bit easier, let’s take the wave function of a particle of which the amplitude only varies along one dimension of space, $x.$

    Here is a representation of a particle’s wave function along one dimension of space (along a ‘straight line’). The $x$-axis represents a position in space. The $y$-axis represents the amplitude of the wave function (which is proportional to the probability of finding the particle in that particular position $x$).

    Figure 8. A representation of a wave function of a free particle. Note that this is not what it actually looks like. For one, an actual wave function exists in complex space, which we didn’t plot here. The goal is to illustrate, not to map accurately. Also note that the free particle has no specific position yet as it’s a free particle!

    It was the eminent French physicist Louis de Broglie(beginfootnote)Many physicists have tried and mispronounced his last name. It should sound like ‘broy’ where the r is produced at the back of the throat, like the French r – a ‘dry’ kind of r. In this interview with him, you can hear the French presenter pronouncing his name (just after 0:16 seconds). It’s not ‘brog-ly’ nor ‘bro-ly’. Thank you.(endfootnote) who formulated the relationship between a particle’s wave function’s wavelength $\lambda$ and its momentum $p.$

    $$\lambda = \frac{h}{p},$$

    where $h$ is the Planck constant. Incidentally, this is the equation better known as De Broglie’s matter wave hypothesis, stating that matter, such as electrons, possess a wave-like characteristic(beginfootnote)Do note that this same equation shows that this wave-like behaviour of large bodies such as our bodies, brains, bowling balls, tennis balls, and animals is completely and utterly negligible as we will demonstrate at the end of this post.(endfootnote). This won him the Nobel Prize, no less.

    If we rewrite this to solve for $p,$ we get

    $$p = \frac{h}{\lambda}.$$

    So, clearly, a wave’s momentum is determined by its wavelength. The smaller the wavelength, the greater the momentum. What is the wavelength? It’s the length between two peaks (or two troughs). The higher the frequency, the smaller the wavelength. Now have a look at Figure 8 again. As you can see, the infinite wave of a free particle has a well-defined wavelength. The logical conclusion is that the momentum is also well-defined. Nevertheless, Figure 8 also shows that the particle’s position is not defined at all!

    Let’s turn this on its head and limit the range of possible positions of our particle. No longer is it a free particle. It is now confined within a finite range of locations.

    Figure 9. Our former free particle’s position is now restrained between $x = 0$ and $x= \pi.$ In other words, $\Delta x$ is now limited to $\pi$ wide. There is no well-defined wavelength as the wave function has different values in different places. It’s there, but not as well-defined as in the wave function in Figure 6.

    What we’ve done in Figure 9 is making $\Delta x$ smaller than it was in Figure 8 (where it was infinitely large). In fact, $\Delta x = \pi$ wide. By the same Fourier transform mechanism as with the time-frequency pair, the complimentary sister of position space $\Delta x$, namely momentum space $\Delta p_x$, will now become less certain.

    To construct a limited wave function such as the one in Figure 9, Fourier analysis shows that you need – again – a bunch of waves at different frequencies in superposition (added on top of each other).

    Figure 10. A Fourier deconstruction of the wave function in Figure 9. Many waves, many frequencies. Hence, the momentum is less well-defined.

    So, when it comes to quanta, Heisenberg’s uncertainty principle states that there’s a fundamental trade-off between information on position and momentum(beginfootnote)Another pair is energy and time. This is interesting in the context of Hawking radiation. We’ll get to that, don’t worry.(endfootnote). This is due to the fact that they are a Fourier transform pair or conjugate variables.

    This also means that if you constrain a particle to a minuscule $\Delta x,$ its wave function will start to contain momenta $\Delta p_x$ all over the place. It will occupy many more velocity possibilities, including the much faster velocities. If you were to subsequently perform a measurement, the probability of finding it moving at higher speeds is now much larger!

    Scale and effect

    At the scale of the big bad world, we never see this effect. If you would confine a bowling ball in a limited space, you will not see its momentum increase dramatically. It won’t suddenly start bouncing up and down. Conversely, if you swoop the bowling ball with considerable momentum, it won’t suddenly start appearing everywhere and nowhere at the same time: its position is still quite clear. You won’t suddenly quantum tunnel through the pins or be rolling on all bowling lanes of the neighbouring players at the same time. If it doesn’t hit a single pin, then that’s not because it’s suddenly in a state of superposition with regard to its possible locations of existence. You’re just not that good.

    You won’t notice any of these quantum effects in your everyday-scaled objects. Only when you’re dealing with particles. Or atoms. However, as soon as the mass increases, it all changes. Why? Partly because Planck’s constant is so darn small(beginfootnote)And because the number of interactions between atoms increase exponentially, causing any quantum effect to disappear due to decoherence.(endfootnote). It’s just $5.273 \times 10^{-35} \text{ J/s},$ remember? That’s small.

    All this knowledge does allow for some fun calculations. For instance, if you were to confine a bowling ball with a mass of $7.2$ kg (16 lb) inside a box where $\Delta x = 22$ cm (8.66 inches), by Heisenberg’s uncertainty principle, the ball’s speed will be $3.283 \times 10^{-35} \text{ m/s}.$ That means that after $965.9$ billion years it might have moved a distance equal to the diameter of a proton. That amount of time is seventy times the age of our current universe. Granted, quantum-mechanical effects aren’t zero, but as you can see (or rather, as one can calculate), on our everyday scale, these effects are quite meaningless.

    Sometimes, weird films such as What the #$*! Do We (K)now!? and What the Bleep!?: Down the Rabbit Hole will want to make you believe such quantum things can happen anyway. They will mention Heisenberg’s uncertainty principle like it is a magical law allowing us to do whatever. I hope that this post has shown that Heisenberg’s uncertainty principle is not about that. Nor does the uncertainty principle itself have its roots in quantum mechanics. It’s basically wave mechanics, the classical stuff, which all first-year undergraduates in physics have to learn in their first or second semester.

    A few months ago, I stumbled across a video showing an Australian senator’s question to the head of the Commonwealth Scientific and Industrial Research Organisation, an Australian federal government agency responsible for scientific research. Clearly, the senator had – shall we say ‘read something about Heisenberg’s uncertainty principle’. During a senate hearing for a legislative committee, the senator questioned if research done in climate change should be taken with precaution as Heisenberg’s uncertainty principle stands in the way of accurate measurements(beginfootnote)He basically sought a ‘scientific’ way to put climate science in doubt – which, apparently, he is not a proponent of. I do not claim to know anything about Australian politics, or even at great depth about climate science, however, when a legislator starts talking quantum physics – well, I do know stuff about that.(endfootnote).

    I suspect this discussion pertained to a study where a satellite uses infrared radiation to perform surface and/or atmospheric remote sensing. He continued to state that as infrared light has lower frequencies than visible light, it’s ‘very difficult’ to understand the properties of infrared radiation based on Heisenberg’s uncertainty principle.

    Many things were going on (wrong) in this one short bit of speaking time of the senator, as is usual when someone hasn’t caught up on quantum physics as much. Which is understandable, but no less gnawing to watch (the link opens a new tab and leads to a short video on Twitter).

    In any case, I genuinely hope that this article contributed at least a sliver of knowledge to educate the electorate of the world, so we can all vote as informed and responsible as possible for the right persons for the right jobs, besides one’s preferred socioeconomic idealism.

    If you should take one thing from this post, it’s that Heisenberg’s uncertainty principle is not about anything spiritual nor does it have anything to do with scientific measurement mistakes: it’s good, old wave mechanics and Fourier analysis taught to undergrads in their first year at university. It works and it works well. It does not lead to science not being able to know things about the universe. In fact, it increased our knowledge of it. In fact, no modern information device would have worked without it. After all, you’re reading this with an electronic device which exists thanks to Fourier, Heisenberg, and De Broglie, among others. All that with a bit of more maths and more physics at the same time.

    Photo Werner Heisenberg by Friedrich Hund, a German physicist who took this photo in Heisenberg’s place of residence, Göttingen, in 1924. It was uploaded to Wikimedia Commons under CC BY 3.0 by Friedrich Hund’s son, Gerhard Hund, a German mathematician, computer scientist, journalist, and chess player. We have used a colour-corrected version by Martin Geisler.

  • Quantum entanglement: the EPR paradox and Bell’s Theorem

    Quantum entanglement: the EPR paradox and Bell’s Theorem


    When the state of a subatomic particle cannot be described by a wave function without taking the state of another subatomic particle into account, we speak of quantum entanglement. It’s the special case where both particles can only be described by one and the same wave function. No longer are they separate entities nor do they have separate wave functions. The astonishing consequence is that performing a measurement on one particle has an immediate effect on the measurement of the other particle, no matter how far apart they are from each other. In this article, the second part of our mini-series on quantum entanglement, we will discuss the EPR paradox which Einstein and colleagues put forward. After that, we will discuss Bell’s Theorem which allowed physicists to test Einstein’s proposal. Was Einstein correct?

    A representation of an electron’s spin – do note that this is not what an electron actually looks like nor is it what its spin looks like. The quantum world is simply too strange to depict accurately using ‘classical’ notions as done here. Here we drew a vague ball-like thing which seemingly spins around, which it isn’t and it doesn’t. But it’s the best we’ve got. Although, the best we’ve got is actually something else: a mathematical expression, the wave function.

    Quick summary

    Firstly, let me give a quick summary of the previous post:

    1. we used the property of spin as a way of distinguishing between the two entangled electrons;

    2. the orientation of an electron’s spin is expressed as spin up (anticlockwise) or spin down (clockwise) along the axis of measurement;

    3. you can arbitrarily choose along which axis you want to measure its spin, in three dimensions;

    4. no matter which axis you choose, the result is always going to be a spin up or spin down (there is no spin-a-bit-to-the-right, for instance);

    5. we are able to entangle particles in such a way that they will either always yield opposite spin or they always yield identical spin; once prepared this way, they will never deviate from this correlation when measured;

    6. we used the opposite-spin entanglement in our example and we will do so again here;

    7. quantum mechanics states that before measurement neither electrons have a specific spin: the wave function contains all possible measurement outcomes, in this case pertaining to both spin up and spin down (which can be characterised as having no definite spin yet)(beginfootnote)Analogously, the double-slit experiment showed that before measurement, particles don’t have a specific location yet.(endfootnote);

    8. as soon as you measure one electron’s spin along a certain axis, the other electron’s spin immediately snaps to the opposite orientation along that same axis, regardless of spatial distance between the two entangled particles(beginfootnote)Or, if their entanglement were prepared in such a way that they always have identical spin, the other electron would then immediately snap to the identical spin orientation along the same axis of measurement.(endfootnote).

    EPR paradox

    Even though Einstein understood quantum mechanics like few others, and while accepting these predictions and results, he didn’t quite like the non-local implications brought forth by quantum entanglement. He didn’t like point 8 of the previous section. There seems to be zero time delay between influencing a particle in Amsterdam (through measuring its spin) and influencing its entangled particle in Boston. It violates a pivotal consequence of Einstein’s theory of special relativity: no signal or piece of information – anything within this universe, really – can exceed the speed light(beginfootnote)In a vacuum.(endfootnote) or else causality would not exist. In other words, if information or signals were able to travel faster than light, an effect could occur before its cause had taken place. To put it mildly, this doesn’t seem to be the universe you and I are living in.

    So, Einstein, Podolsky, and Rosen (EPR) hypothesised that something else, something secretive was going on in nature – well out of sight for theoretical and experimental physicists. Quantum mechanics as it was known then had to be incomplete. Obviously, they acknowledged its successes, but when it came to quantum entanglement, they asserted something was missing in the theory of describing nature through wave functions.

    To solve for the seemingly faster-than-light signal, they proposed that what really was going on was that the particles have always been in a specific state. When the electron pair were separated from each other, they have always had either spin up or spin down from the start from the moment of their creation.

    Suppose, a pair of gloves were made. Like all pairs of gloves, they always were each other’s opposite with respect to ‘handedness’(beginfootnote)‘Handedness’ in this context is a form of the more generalised term chirality.(endfootnote). One has always been left-handed, the other has always been right-handed. And if the first one happened to be right-handed, then the other was left-handed. (Or else you’re holding a glove from another pair.)

    Suppose, the machine which had made the pair put each glove in a separate box. We can’t see which glove went in which box until we open the box. The boxes were sent to Amsterdam and Boston. The experimental physicists then open the box in Amsterdam: it’s the right-handed one! And so, we now instantly know, the one in Boston is left-handed. No magic, no non-locality, no lightspeed-breaking shenanigans.

    This is what Einstein and friends said was happening in the case of electrons. An electron pair always had specific spins to start with. It’s only in Amsterdam and Boston that we ‘open the box’ aka measure their spin. It’s only logical now that as soon as you know which spin the Amsterdam electron has, you immediately know which spin the Boston electron has.

    So, said Einstein, non-locality is an illusion. It’s all just normal local laws of nature and a bit of logical thinking. For one, spin orientation is merely hidden from us and not principally uncertain. Secondly, there’s no spooky action at a distance[1], as he famously described it(beginfootnote)In German, he wrote ‘spukhafte Fernwirkung'[1].(endfootnote).

    In everyday parlance, physicists call this a local version of the ‘hidden variables’ theory. ‘Hidden variables’ pertain to the stuff that we can’t see yet (such as spin orientation or other variables influencing this) because our quantum mechanical description (the wave function) is incomplete, however, they are there, they do exist – they do not not exist yet, according to the hidden variables theory.

    Bell’s inequalities

    Unfortunately, Albert Einstein passed away in 1955. And Niels Bohr, the other great physicist with whom he used to debate the fundamental nature of quantum mechanics passed away in 1962. In both cases too soon for them to be able to read John Stuart Bell’s 1964 paper called ‘On the Einstein Podolsky Rosen Paradox'[2]. Bell realised that Einstein’s proposal was in principle testable. It yielded a clear prediction, called Bell’s inequality.

    At this point, we must note that over the years, more than one Bell’s inequalities have been put forward by physicists(beginfootnote)Besides his original inequality, there’s the much-used CHSH-inequality, for instance.(endfootnote). To explain Bell’s inequality, we will apply a version of David Mermin’s original version as mentioned in his fantastic Boojums All the Way Through: Communicating Science in a Prosaic Age[3].

    Recall from point 3 before that we can measure an electron’s spin orientation along any axis. We’re going to be measuring along three axes. These axes will be at an angle of 120° relative to each other.

    The first axis will be the spin orientation along the vertical axis, which we will denote with the following symbols for spin up and spin down:

    $$\uparrow \downarrow$$

    The spin orientations up and down will also be measured along this second axis:

    $$\nwarrow \searrow$$

    And the spin orientations along the third axis will be denoted by:

    $$\nearrow \swarrow$$

    So, imagine two entangled electrons being separated in space from each other. The usual quantum-mechanical description of each electron is that they are in a superposition of spins up and spins down for all three axes.

    Except, Einstein says, no, no, not really: hidden behind the ‘veil of superposition’ they are in fact already in definite, specific spin orientations for each of the three axes. We just don’t yet know which until we measure them!

    He says, the electron in Amsterdam may already be in the specific spin states as follows:

    $$\left( \uparrow \searrow \swarrow \right)_A$$

    So, along axis 1 it’s spin up, along axis 2 it’s spin down, and along axis 3 it’s also spin down.

    Einstein continues and says that the entangled electron in Boston has to already be in the opposite states:

    $$\left( \downarrow \nwarrow \nearrow \right)_B$$

    And so, Einstein concludes, as soon as you actually perform a measurement in Amsterdam along the first axis, of course, you get the opposite spin in Boston. Only logical!

    Bell’s insight was that if you would work out this entire argument for all possible combinations, you could actually get a prediction of a ratio of outcomes. Here’s how that goes.

    First of all, if you measure along axis 1 in Amsterdam, that doesn’t mean you have to measure along that same axis in Boston. You could just choose to measure along axis 3. So, with the two examples above, your results would simply be that in Amsterdam you get spin up and in Boston you also get spin up:

    $$\left( \uparrow \right)_A \text{ and } \left( \nearrow \right)_B$$

    Bell then argued, if you would count the number of times you would get the combinations up-up, down-down, and of course up-down and down-up like this, you should get ratios of these combinations which should match experiment. If, however, these ratios don’t appear in the experiments, then Einstein’s hypothesis is incorrect. In that case, something entirely different is going on. The electrons were not already in a specific state, which in turn means that the non-local measurement effect in quantum entanglement does exist!

    Bell’s theorem

    So, let’s put them all together. Let’s first take our example above:

    $$\left( \uparrow \searrow \swarrow \right)_A \text{ and } \left( \downarrow \nwarrow \nearrow \right)_B$$

    If you measure along axis 1 in Amsterdam and along axis 1 in Boston you get spin up, spin down. If you measure along axis 1 in Amsterdam and along 2 in Boston, you get spin up, spin up. And so on, and so forth! We’ve put it in a little table:

    Here you can see all the possible combinations of measurement outcomes along the three possible axes of the electrons in Amsterdam (A) and Boston (B). We used U for spin up and D for spin down.

    Bell then says that if Einstein was correct, and the states of the spin orientations along these three axes were already there, then these are the expected outcomes.

    Let’s focus on the number of UD or DU combinations, in other words, let’s focus on the number of times we find the opposite spin orientations, irrespective of the axes along which they are measured. We’ve marked them yellow.

    Exactly five out nine times you will find the opposite spin directions.

    Let’s check for other spin combinations. Suppose, the electron in Amsterdam is secretly in the following spin states, $\left( \downarrow \nwarrow \swarrow \right)_A$, and the electron in Boston is then the opposite, $\left( \uparrow \searrow \nearrow \right)_B$. If we count again the number of times the measurement outcome of opposite spins, we get, again, five out of nine.

    Okay, I think you can imagine where this is going. We’re not going to go by all the tables, but I do want to do one more, just for fun. Suppose, the one in Amsterdam is all spin down, $\left( \downarrow \searrow \swarrow \right)_A$, and, obviously, the Boston one is its opposite, $\left( \uparrow \nwarrow \nearrow \right)_B$. In that case, we would get opposite spins in nine out of nine times.

    And so, this particular Bell inequality states that the probability (P) of finding opposite spins along all three axes is at least $\frac{5}{9}$ or 55% (and at most 1 or 100%). In other words, $P(\text{opposite}) \geq \frac{5}{9}$. If this inequality were violated by experiment, the underlying theory will have been proven to be incorrect.

    Experimental outcomes

    Over the past thirty years, many experiments were carried out to test multiple versions of Bell’s inequality. Usually, these tests involved photons rather than electrons and pertained to measurement of polarisation rather than spin.

    Freedman and Clauser did the first Bell test. They used a version of the so-called CH74 inequality[4].

    The most well-known test was performed by Alain Aspect and colleagues. As Bell had originally suggested, they were able to have the two measurement devices randomly select the method of measurement before the entangled photons had arrived[5].

    In all tests, all versions of Bell’s inequalities were violated. Instead, the statistical outcome was congruent with the predictions of quantum mechanics. The conclusion has to be that Einstein’s local hidden variable theory was incorrect. There is nothing local about measuring entangled particles.

    In our particular inequality, the result was that the occurrence of opposite spins turned out to be exactly 50%, not 55%.

    Conclusions

    Let’s summarise what we have established over the course of the last two posts, including this one.

    In quantum mechanics, particles which have not been measured yet don’t have a definite, specific state. Instead, they are best described by a wave function which incorporates all the possible future states it can snap into once measured.

    When a particle can only be described in tandem with another particle, i.e. both particles can only be described by one and the same wave function, they are maximally quantum entangled(beginfootnote)In practice, in the real world, particles aren’t maximally entangled like the way we can prepare them in the laboratory. The world is too messy for those ‘pure states of entanglement’ to exist for any significant amount of time. There are simply too many particles around to not interact with any other particle. Every particle will invariable interact with thousands of trillions of other particles and so any previous entanglement will quickly decohere into either a very weak version of the original entanglement or simply to zero entanglement. Every interaction represents a measurement. Since our brains are too large and consist of thousands of trillions of particles, they will never be in a pure state of superposition nor entanglement. Not to mention our much larger body, which will never be in any sort of quantum state. It is statistically so unlikely that you’d have to become as old as $(10^{100})^{100}$ times the age of our current universe to witness such an event. And that number was a metaphorical one. It’s much larger.(endfootnote).

    If their entanglement entails their spins will always correlate in a certain way – be it identical spins or opposite spins – a measurement on one particle, causing it to snap into one of the possible, specific, definite states, has immediate effect on the state of the other particle: it instantly snaps out of its wave function haze into a correlating, specific, definite state.

    Einstein didn’t like this as this would imply some kind of information was somehow transported beyond the speed of light from one particle to the other.

    He postulated that particles have always been in a specific, definite state to begin with. The only reason we don’t know which is because we haven’t measured it yet. There is no ‘snapping out of the haze’ going on.

    John Bell showed that Einstein’s hypothesis can be tested. If you would perform many, many measurements of many, many maximally entangled particles, eventually, the occurrences of the variety of correlated states should show up in a certain ratio, an inequality, as it happens.

    Experiments showed they do not. Instead, the ratio is exactly according to the predictions of quantum mechanics.

    This demonstrated that particles indeed snap out of their haze upon measurement and not that particles had always been in a hidden but definite state.

    And if that is true, then non-locality has to be true – there is no other way the other particle snaps into the correct, correlated state.

    Nobody knows how this happens. Certain non-local but still hidden-variables hypotheses have been proposed. One of the more famous versions is called the ER=EPR conjecture by Juan Maldacena and Leonard Susskind. Perhaps we’ll dive into that later on.

    Einstein’s aversion to this ‘particles have no definite state until measured upon’ made him utter his famous complaint, ‘God does not play dice’.

    Unfortunately, he was wrong here on two occasions. God(beginfootnote)We are using the word ‘God’ in a purely metaphorical way. This does not pertain to any specific religious entity as revered by many in a variety of societies in human culture.(endfootnote) does play dice. Moreover, He throws them where we can’t see them. Even God seems to be bound by Heisenberg’s Uncertainty Principle. But that’s a subject for another bit of maths and physics.


    [1] Einstein, A., Podolsky, B. and Rosen, N. (1935) “Can Quantum-Mechanical Description of Physical Reality Be Considered Complete?,” Physical Review, 47(10), pp. 777–780. doi: 10.1103/PhysRev.47.777.

    [2] Bell, J. S. (1964) “On the Einstein Podolsky Rosen Paradox,” Physics Physique Fizika, 1(3), pp. 195–200. doi: 10.1103/PhysicsPhysiqueFizika.1.195.

    [3] Mermin, N. D. (1990) Boojums all the way through : communicating science in a prosaic age. Cambridge England: Cambridge University Press.

    [4] Fry, E. S. and Thompson, R. C. (1976) “Experimental Test of Local Hidden-Variable Theories,” Physical Review Letters, 37(8), pp. 465–468. doi: 10.1103/PhysRevLett.37.465.

    [5] Aspect, A., Dalibard, J. and Roger Gérard (1982) “Experimental Test of Bell’s Inequalities Using Time-Varying Analyzers,” Physical Review Letters, 49(25), pp. 1804–1807. doi: 10.1103/PhysRevLett.49.1804.

    Featured image: Portrait of theoretical physicist John Bell at CERN, June 1982 (CERN, CC BY 4.0)

  • Quantum entanglement: non-locality and the state of a two-particle system

    Quantum entanglement: non-locality and the state of a two-particle system


    To this day, quantum entanglement and its effects are phenomena which still leave physicists scratching their heads when trying to get a deeper understanding of what is actually happening. This series on quantum entanglement is going to be a two-parter. In this post, we will discuss what is meant by locality and non-locality and what quantum entanglement is. The term quantum entanglement has been used in many instances of popular culture pertaining to spirituality, healing, and a flurry of new age approaches to human consciousness. This is not the kind of ‘quantum entanglement’ we will discuss here. We will purely look at the physics of it, its original and proper meaning. We will study the state of a two-particle system. In the next post, we will discuss what Einstein and his friends proposed, what Bell wrote, and whether Einstein was right. And then there are also exciting caveats which we will explore.

    The basics

    Let’s go over the basics one more time. ‘Particles’ aren’t particles in the classical sense at all – they’re absolutely not like tiny balls or pellets. They are best described by the wave function, a mathematical expression containing all possible states the particle can be in. This pertains to its energy levels, its positions or a number of other properties it can have.

    As long as no measurements have been performed on it, the particle has no definite state or states. It displays wave-like behaviour like being caught in a haze of all possible states. However, as soon as you measure it, the particle will snap out of its haze and it will appear to be a particle, an actual particle in the classical sense, with a definite state.

    Note that ‘the state of an electron’ can refer to a particle with no definite set of states when no measurement was performed. The state of an electron is then best described by the wave function, which contains all possible definite states upon measurement.

    Hereafter, ‘wave function’ and ‘state’ are used interchangeably.

    In This is not an atom, the wave function is discussed. In The double-slit experiment, the wave-like and the particle-like behaviours are showcased.

    Locality vs non-locality

    Isaac Newton knew he had a problem when he formulated his theory of gravity. While it beautifully described the extent to which two masses exert gravitational forces upon each other, his theory didn’t explain how they did that. He didn’t like the conclusion that the gravitational influence between Earth and the Moon seemed to spookily operate at a distance through the vacuum. He wrote it was ‘so great an Absurdity that I believe no Man who has in philosophical Matters a competent Faculty of thinking can ever fall into it’. He famously stated to leave this unsolved mystery to ‘the Consideration of my readers’[1].

    In other words, Newton wasn’t big on non-locality. And yet, his own theory did entail an invisible force operating over vast distances through the vacuum. Moreover, it seemed to be an instantaneous effect: if the Sun were to suddenly disappear, then Earth would be flung off its trajectory immediately. Of course, today, we know that nothing can travel faster than light, so the gravitational changes of the Sun would take about eight minutes to ‘reach’ Earth.

    The following years, physical phenomena such as magnetism and electricity proved, in fact, to be very local indeed. It became clear there is always an indirect way through which one object is able to influence another object at a distance. What is meant with locality? Here’s the mechanism: an object interacts with its immediate environment, a field embedded within the three-dimensional space we live in, i.e. the electromagnetic field, which then passes on that ripple of disturbance onto the other object. In terms of ‘fields’, one could say that at one particular location the field’s value is changed by some object. That value change then changes the values of the field in the direct vicinity, which then change the values in their vicinity, and so on. It’s a bit like ‘the wave’ done by thousands of sports fans in a stadium. Or like falling dominoes. Every change is ever local and the propagation of that change through space is limited to the speed of light.

    Tumbling telephone boxes are definitely a ‘local phenomenon’. The sculpture Out of Order by David Mach is situated in Kingston upon Thames (UK). Photo by 272447.

    Many years later, Einstein replaced Newton’s theory with his own theory of gravity, General Relativity (GR). It showed that Newton’s intuition was correct. Gravity couldn’t be non-local and Einstein showed it isn’t. In GR, space and time itself are the stretchy substance through which gravitational disturbances propagate at the speed of light towards the other object. When a mass curves or disturbs spacetime around it, that curvature or disturbance then ripples through the universe, on its way to influence other objects. In fact, on 11 February 2016, a large collaboration of incredibly talented scientists physically measured these gravitational ripples in spacetime as predicted by Einstein in 1916. It won three key figures the Nobel Prize.

    And so, it seems there is no spooky influence at a distance in physics. Even still to this day, in modern quantum physics, our best understanding and most successful theory is that quantum fields pervade our universe, forming the mediums through which forces are propagated, limited by the speed of light.

    Non-locality entails a change in one patch of space instantaneously influencing another patch of space irrespective of their distance. Locality entails the propagation of change through space by influencing only neighbouring patches of space at a maximum of the speed of light.

    Spin

    Electrons have several properties. One of the more obvious is (negative) charge. The Stern-Gerlach experiments showed that they possess another property which was given the name spin angular momentum or simply spin for short, for lack of a better term as electrons aren’t exactly like spinning balls.

    Nevertheless, as it stands, electrons have an intrinsic spin, which cannot in any sensible way be described like a classical-mechanical rotation. Like with any object in three-dimensional space, you can measure its spin along any angle within 360 degrees in three dimensions. With respect to whichever axis you choose, they can only ever spin clockwise or anticlockwise(beginfootnote)Yes, this does sound like there is an actual rotation around an axis in the classical sense. And maybe, in some deep sense, there is after all, however, this deserves a post of its own, so suffice to say for now, our language is simply too limited to avoid using classical terms for quantum mechanical phenomena, misleadingly.(endfootnote). The latter is called spin up and the former spin down, according to the right-hand rule.

    If electrons were like tiny, fluffy balls such as displayed here, you could picture their spin as an anticlockwise or clockwise rotation about the axis of measurement. Using the right-hand rule, we can designate this spin-up or spin-down. Of course, in three dimensions, any axis of measurement at any angle can be chosen with respect to which it will be found spinning. Disclaimer: this classical-mechanical illustration does not portray actual electrons nor actual quantum mechanical spins. But it’s perhaps useful as a simile. (Illustration by KJ Runia)

    Symbols

    As we take our readers seriously, we’ll take this opportunity to introduce a few mathematical symbols which will prove to come in handy at later stages of this series.

    Let’s use the symbol $\lvert A \rangle$ to denote the state of the electron in Amsterdam with respect to its spin. As long as we haven’t performed any measurements on the electron, it has no definite state. However, upon measurement, its spin with respect to the vertical axis of measurement is ever either spin up or down. Let’s write these two possible measurement outcomes as $\lvert\uparrow\rangle_A$ or $\lvert\downarrow\rangle_A$.

    Likewise, if the state of an electron in Boston $\lvert B \rangle$ is spin up or spin down, we write $\lvert\uparrow\rangle_B$ or $\lvert\downarrow\rangle_B$.

    Assuming the state of the electron in Amsterdam hasn’t been measured yet, we can express this (with respect to spin) as a combination of both spin states:

    $$\lvert A \rangle = \alpha \lvert\uparrow\rangle_A + \beta \lvert\downarrow\rangle_A .$$

    This is why physicists often poetically say that the unmeasured particle is in a state of both spins at the same time while it’s more accurate to say it has no definite state. Mathematically, its state is an amalgam of all possible, linearly superposed (added together), algebraic solutions to the Schrödinger equation, hence, it’s said to be in quantum superposition.

    What’s that $\alpha$ and $\beta$, you ask? Well, they’re numbers of probability we need to find in order to complete our expression. The Born rule states that if we square the (modulus of the) wave function (the state), we will get the probability (density) of either possible outcome after measurement. Now, experiments have shown that either outcome, spin up or spin down, $\lvert\uparrow\rangle_A$ or $\lvert\downarrow\rangle_A$, appears in 50% of the total number of measurements. In other words, the probability of measuring either spin state is exactly $\frac{1}{2}$. So, if we put $\alpha=\beta=\frac{1}{\sqrt{2}}$, then $\lvert\alpha\rvert^2 = \lvert\beta\rvert^2 = \frac{1}{2}$. After all, $(\frac{1}{\sqrt{2}})^2 = \frac{1}{2}$, which is exactly what we want. So, the state (wave function) of our Amsterdam electron with respect to spin can be represented by

    $$\lvert A \rangle = \frac{1}{\sqrt{2}} \lvert \uparrow\rangle_A +\frac{1}{\sqrt{2}} \lvert \downarrow\rangle_A .$$

    Similarly, the state of the electron in Boston with respect to spin is then represented by

    $$\lvert B \rangle = \frac{1}{\sqrt{2}} \lvert \uparrow\rangle_B +\frac{1}{\sqrt{2}} \lvert \downarrow\rangle_B .$$

    What you need to take from this is the following: the state of an electron before measurement is the sum of all possible states (multiplied by a probability factor, in this case $\frac{1}{\sqrt{2}}$).

    In the case of spin as measured along the vertical axis, the state of the electron is the sum of two possible states, spin up $\lvert \uparrow \rangle$ or spin down $\lvert \downarrow \rangle$.

    Note that there are other possibilities: we could measure the spin along a horizontal axis. We could represent this with spin left $\lvert \leftarrow \rangle$ or spin right $\lvert \rightarrow \rangle$. Or we could measure the spin at angles of +120 or -120 degrees from the vertical axis, which we might represent as $\lvert \nwarrow \rangle$ and $\lvert \searrow \rangle$ or $\lvert \nearrow \rangle$ and $\lvert \swarrow \rangle$. We will get to that in the discussion of Bell’s Theorem in the next post.

    Quantum entanglement

    So, what is quantum entanglement? Recall that the most complete description of a particle is the wave function. This has always been about a free, single particle, not interacting with anything. In the case of quantum entanglement, however, this doesn’t fly anymore.

    When the state of a particle can no longer be described without a description of the state of another particle, those two particles are said to be quantum entangled. No longer can we describe either particle by one wave function each. They can only be described as a two-particle system by one and the same wave function.

    This has an astonishing consequence. Suppose our two electrons become entangled in such a way that they always have opposite spins(beginfootnote)Producing spin-entangled electrons is difficult but clever experimental physicists have their ways.(endfootnote). So, if one has ‘spin up’, $\lvert \uparrow \rangle$, the other always has ‘spin down’, $\lvert \downarrow \rangle$, or vice versa(beginfootnote)It’s also possible to have them correlate such that they have identical spin, but for our example, let’s not.(endfootnote). So, we now have one system with two particles who always have opposite spins, which means that the total spin of our system is 0, zero. Let’s denote the total spin of our system with $\lvert S \rangle$.

    Before our experiment takes place, they are both separated. One is staying in a laboratory in Amsterdam. The other is transported to Boston. Since no measurement has taken place on either particle, they are in a superposition according to the one wave function. They haven’t an exact location (although one is very likely to be somewhere in Amsterdam at the moment of measurement and, likewise, the other in Boston), their energy levels are all over the place, and their spin isn’t either spin up or spin down along this or that axis.

    We can represent this whole situation with respect to spins as follows:

    $$\lvert S \rangle = \dfrac{1}{\sqrt{2}} \left( \lvert \uparrow \rangle_A \lvert \downarrow \rangle_B – \lvert \downarrow \rangle_A \lvert \uparrow \rangle_B \right) .$$

    When you’re looking carefully at the expression above, you can see that the state of the total spin $\lvert S \rangle$ of our two-particle system is a combination of two situations: the electron in Amsterdam is spin up and so the electron in Boston is spin down or the electron in Amsterdam is spin down and the electron in Boston is spin up. They need to be subtracted from each other because the total spin equals 0, remember? Hence, the minus sign. Lastly, both states are multiplied by the fraction $\frac{1}{\sqrt{2}}$ because both states have a 50% chance of occurring (which you get if you square the whole thing).

    And so, what does this mean? As soon as you perform measurements on the one in Amsterdam, and you find it has spin up, the other electron in Boston immediately has spin down along that particular axis upon measurement, even though the probability before measurement was still 50%! How does the electron in Boston ‘know’ what the measurement result in Amsterdam was? En how does it know this so fast? Faster than the speed of light! Besides this, turns out, you’ll always get a definite spin from the other particle opposite to the one you measured first. As soon as the measurement in Amsterdam took place, the measurement outcome in Boston being the opposite result is always 100% all of a sudden! (Or the other way around.) There are never any exceptions!

    In other words, as soon as you do the measurement, the mathematical description changes from

    $$\lvert S \rangle = \dfrac{1}{\sqrt{2}} \left( \lvert \uparrow \rangle_A \lvert \downarrow \rangle_B – \lvert \downarrow \rangle_A \lvert \uparrow \rangle_B \right) ,$$

    to either

    $$\lvert S \rangle = \lvert \uparrow \rangle_A \lvert \downarrow \rangle_B ,$$

    meaning, the state of the total spin equals the one in Amsterdam being spin up and the one in Boston being spin down, or, vice versa:

    $$\lvert S \rangle = \lvert \downarrow \rangle_A \lvert \uparrow \rangle_B .$$

    And here’s the astonishing part: this will always work this way, no matter how great the physical distance between the two particles. Locality out the window. Welcome back, non-locality.

    Einstein accepted this prediction in quantum mechanics as being correct. However, he didn’t like it. How did the other particle instantly ‘know’ which spin to exhibit when Einstein’s fantastically successful theories of relativity relied on the universal law that nothing can exceed the speed of light? He accepted the theory but he concluded it wasn’t complete. There had to be some sort of hidden mechanism which they had overlooked.

    We will discuss Einstein’s attempt at saving the principle of locality and the universal speed limit in the next post. As well as John Bell’s and Alain Aspect’s subsequent work. For now, the question of whether Einstein was right, we will ‘leave up to the Consideration of our readers.’


    [1] Newton, I. (1756) Four Letters from Sir Isaac Newton to Doctor Bentley: Containing Some Arguments in Proof of a Deity [Online]. Available here. (Accessed: 14 May 2020)

    Featured image by KJ Runia

  • Lab centrifuges and prime numbers

    Lab centrifuges and prime numbers


    When micro- or molecular biologists do research on viruses, bacteria, fungi, human or animal cells, one of the many instruments they will use is a laboratory centrifuge. This equipment allows them to separate substances contained within a test tube. This way scientists are able to obtain, for instance, purified enveloped viruses, such as the novel coronavirus, SARS-CoV-2. Or they can isolate nucleic acids, such as DNA.

    Often, the rotor of the machine rotates at incredible speeds. It is vital that the test tubes have been placed in a perfectly balanced way. If not, the machine might break down and potentially dangerous glass shards and substances might be flinging about(beginfootnote)Although sensors may be installed to prevent the machine from operating in case of force imbalance. See also the Final remarks down below.(endfootnote).

    Fortunately, there is a nifty way to calculate whether you can – in principle – place a certain number of test tubes in an evenly balanced way. To crack the code, we will use my favourite type of number: the prime numbers. Fun fact: this funky little trick wasn’t proven until fairly recently in 2010.


    NEW: Listen to the audio |


    The set-up

    Before we begin, we assume that the mass of each test tube, including their contents, is equal. Also, I would like to remark that, of course, we could do this the physics way, using angular velocity and torque and all that, but in this case, we’re going to be all mathy about it, or specifically, in a way, number-theoretical.

    Suppose, the machine can hold eight test tubes. Eight holes are positioned in a circle on the rotor bit of the machine.

    If we have just one test tube, there’s no way we can make it balanced. That much is clear. If we have two test tubes, however, no problem. They can be balanced easily. Just put one on either side precisely opposite each other. Three test tubes? Hm. I don’t see how. Whatever arrangement we try, it’s always going to be asymmetrical. What if you have four test tubes? Well, this is easy enough. Make it symmetric, like a square.

    Okay, so what about five test tubes? Well, that’s just the same as when we had the inverse of this, with three test tubes! That couldn’t be done, so, this can’t be done either.

    Six? Yeah, of course, we can do that. It’s just the same as having two test tubes, it’s just the inverse! Three on one side and three on the other side. Now you have two open spots on either side. Perfectly symmetrical, just like the inverse situation, where you had two test tubes and six open spots.

    Seven? No. You will have guessed it by now. Having seven test tubes is exactly the same as having just one test tube in a rotor with eight spots.

    And eight, well, of course, we can do eight. It’s also the exact same as having no test tubes at all. So, yes, that’s balanced.

    Do you see a pattern here? You might. Notice how the number of occupied spots and empty spots always complement each other.

    Prime factorization

    Just for clarity’s sake, I’m going to call whole numbers integers since that’s what they’re called in mathematics.

    So, I’m assuming we all know what a prime number is: an integer greater than 1 which cannot be formed by multiplying two smaller integers. In high school or even in primary school, you may have been taught that prime numbers are numbers which can only be divided by 1 or by itself (not including 1). So, prime numbers are 2, 3, 5, 7, 11, 13, 17 and so on.

    Prime factorization is writing down any non-prime integer as a multiplication of two or more prime numbers. The fundamental theorem of arithmetic states that any integer is either itself a prime number or can be written as a product of prime numbers. This is one of the reasons why they’re my favourite. Primes are the building blocks of any integer.

    So, for instance, we take the number 15. This number can be written as $ 15 = 3 \times 5 $. Or take 279. We can write $ 279 = 3 \times 3 \times 31 = 3^2 \times 31 $. Let’s take 16. This number can be written down as $ 16 = 2 \times 2 \times 2 \times 2= 2^4 $.

    As you can see, prime factorization is pulling apart a non-prime number into a product of prime numbers. We call the latter prime factors. 

    So, that’s what that is. One of the many applications of prime factorization is finding the greatest common divisor between two integers, for example. Or encrypting (and decrypting) secret files and messages. Here, we’re going to use it for calculating whether test tubes can be arranged in a balanced way.

    The trick

    Suppose, your machine has $n$ spots available. Suppose, $k$ is the number of test tubes. The number of empty spots is $n-k$. Here’s the trick.

    Determine the prime factors of $n$. If (and only if) $k$ can be written as a sum of these prime factors and the number of empty spots $n-k$ can be written as a sum of these prime factors, you can in principle balance the rotor.

    The mathematics

    It’s too technical to discuss at length the proof given by Gary Sivek in his 2010 paper (or here). However, the gist for the more mathematically inclined is available by clicking ‘expand’. You may skip this paragraph if this is (understandably) still too technical.

    Expand

    Striving to obtain an $n$-th cyclotomic polynomial (or prime polynomial), we obtain a series of complex numbers $z^n$ which satisfy $z^n = 1$, all being $n$-th roots of unity where $n$ is the number of total spots on the centrifuge. We then map the test tubes onto the roots of unity in a non-overlapping way. As is well-known, the values of $z \in \mathbb{C}$ are given by $e^{\frac{2\pi i}{n} k}$, where $1 \leqslant k \leqslant n$.

    So, now we have $k$ roots of unity among the $n$-th roots of unity representing the occupied spots in the centrifuge.

    Sivek proved, using Leung’s and Lam’s Theorem, that if (and only if) the sum of the $n$-powered $k$ roots of unity and the sum of the $n$-powered $n-k$ roots ‘vanish’, i.e. are equal to zero (using good-old de Moivre’s formula, if you remember from your very first semester at uni), as long as $n \geqslant 2$ and $1 \leqslant k l\eqslant n-1 $, then balancing is a fact (where $k=0$ and $k=n$ were regarded to be trivial cases for obvious reasons).

    As you can see, no classical mechanics required.

    An example with eight roots of unity in the complex plane

    Obvious examples

    Suppose, we take our centrifuge which was capable of handling 8 test tubes. We have 6 test tubes. First thing we do is calculate which prime factors the number 8 has. We know this, it’s all 2s. So, the only prime factor of 8 is 2. We can write the number of test tubes, 6, as a sum of this prime factor 2: $6 = 2 + 2 + 2$. The number of empty spots, that’s $8-6 = 2$, is the prime factor itself! So, yes, if you have 6 test tubes, you can balance the machine.

    Let’s take 7 test tubes. Can this be written as a sum of the prime factors of 8? No, it can’t. Well, that’s it then. We cannot arrange the test tubes in such a way that it’ll be balanced out.

    A counter-intuitive example

    Suppose, our centrifuge is capable of handling 12 test tubes in total. We only have 7 test tubes. Hm. Surely, we can imagine 6 test tubes working, but can we make a balanced arrangement with 7 test tubes?

    Let’s first do some prime factorization with 12. So, $ 12 = 2 times 2 times 3 = 2^2 times 3 $. In other words, the prime factors of 12 are 2 and 3.

    Now, can we write 7 as a sum of these prime factors? Yes, we can: $7 = 2 + 2 + 3$. Okay, so far, so good. Can we write the number of empty spots as a sum of these prime factors? Well, $12-7 = 5$. And yes, we can also write 5 as a sum of 2s and 3s: $5 = 2 + 3$.

    So, yes, we can balance 7 test tubes in a rotor with 12 spots! It’s likely this outcome wasn’t immediately apparent to you. If you were to see or draw a depiction and a working out of the arrangement yourself, however, I think it’ll become clear how this would work. Bonus points if you can draw a balanced configuration for 5 test tubes. Because you should know by now, you can.

    Bonus trick

    The beauty of it all is that all of the above does give us another quick way to assess whether we can balance the centrifuge. I’m going to be honest with you: it may be the easiest. If you can express the number of test tubes as the sum of two numbers of which you already know you can balance the rotor, then you can balance the rotor. Heh.

    Final remarks

    In real life, most machines have sensors to prevent force imbalances from taking over. The rotors have markings so that users won’t have to think about where to place the test tubes. Besides, in a university lab, you would simply make sure you prepare the number of samples which make a balancing act trivial. Moreover, many rotors contain three compartments containing sets of test tubes. This makes adjusting for mass variability much easier. And some machines, in hospital labs, for instance, have fully automated robots doing the heavy lifting.

    Therefore, the reason for why this type of mathematics is done, isn’t so much for the applicability as it is for the joy of exploring deep connections such as between prime numbers and complex geometry, if you will. It’s first and foremost a fun and fruitful exercise of human exploration of the lands of number theory, algebraic geometry, and finite fields, on the continent that is pure mathematics.

    Featured image by Michail Tzortzatos under CC BY-SA 4.0
    Spinning rotor by user musicalwoods under CC BY-SA 2.0

  • Finding the normal force in planar non-uniform circular motion using polar coordinates

    Finding the normal force in planar non-uniform circular motion using polar coordinates


    In this post, we will derive an expression for the normal force on a uniform mass which is in planar non-uniform circular motion using polar coordinates. Finding this expression is enormously useful to calculate under which circumstances a mass would be slung off its orbital path. Of course, there are numerous situations for which we should be able find the normal force. Here, we will look at a system as shown in Figure 1. Sometimes, obtaining an expression in terms of the variables given is not straightforward. You will find a useful trick in step 7 to arrive at an expression in terms of a simple $\theta$ instead of its secondary-order derivative $\ddot\theta$ which we initially obtain.

    This could be seen as an undergraduate-physics-level post. Download this article


    Notation

    We will apply Newton’s notation (the dot notation) whenever possible as this is the most compact form. For instance, if $\mathbf{x}$ is a vector, then its first-order and its second-order derivative with respect to time $t$ are denoted by

    \[ \dot{\mathbf{x}}\text{ and }\ddot{\mathbf{x}}, \]

    respectively. Where needed, in order to state explicitly that we are dealing with a time-derivative and to help in solving a time-integral for example, we will use Leibniz’s notation, i.e.

    \[ \frac{\text{d}\mathbf{x}}{\text{d}t}\text{ and }\frac{\text{d}^2\mathbf{x}}{\text{d}t^2}. \]

    Assignment

    Look at the system as sketched in Figure 1. Imagine we stand in front of this system. Mass $m$ is attached to a model string. At $t=0$, it rests at level with the centre of the cylinder with radius $R$ with the string draped over the top. A constant force $\mathbf{P}$ pulls the string downwards. At a later time $t$, mass $m$ has slid over the top with a coefficient of friction $\mu$. Let $\theta$ denote the angle between its initial and its current position, subtended at the centre of the cylinder. Calculate the normal force on $m$, and, hence, proof that the radius of the cylinder is irrelevant.

    Figure 1. The system

    Step 1. Force diagrams and unit vectors

    It is essential to draw force diagrams and unit vectors to define the acting forces and parameters. We choose the unit vectors to be the radial and the tangential vectors. This makes calculating most forces a lot easier. This is done in Figure 2.

    Figure 2. Force diagram and unit vectors at time $t>0$

    We identify the following forces on $m$:

    • $\mathbf{P}$ is the vector denoting the constant force pulling the model string,
    • $\mathbf{N}$ is the vector denoting the normal force acted on $m$ by the cylinder,
    • $\mathbf{F}$ is the vector denoting the frictional force,
    • $\mathbf{W}$ is the vector denoting the weight of $m$ as a result of the gravitational field of whatever planet the system is located,
    • $\mathbf{e}_r$ is the radial unit vector,
    • $\mathbf{e}_\theta$ is the tangential unit vector.

    Step 2. Apply Newton’s second law

    As this is a dynamical system, where $m$ is in non-uniform circular motion, we apply Newton’s second law, more specifically in the following form:

    \begin{equation}
    \sum\mathbf{F} = m\ddot{\mathbf{r}},
    \end{equation}

    where $\ddot{\mathbf{r}}$ is the rate of change of the rate of change over time, that is, the second time-derivative of the displacement vector $\mathbf{r}$ of mass $m$. We can now easily identify the constituents of the vector sum as we did that already in Step 1. And so, equation (1) becomes

    \begin{equation}
    m\ddot{\mathbf{r}} = \mathbf{P} + \mathbf{N} + \mathbf{F} + \mathbf{W}.
    \end{equation}

    Step 3. Rewrite the forces in terms of their magnitudes and unit vectors

    As pulling force $\mathbf{P}$ with magnitude $|\mathbf{P}|$ acts in the direction of tangential unit vector $\mathbf{e}_\theta$, we can write for $\mathbf{P}$:

    \begin{equation}
    \mathbf{P} = |\mathbf{P}|\mathbf{e}_\theta.
    \end{equation}

    Since we don’t have any other information regarding this force, we leave it at that.

    Normal force $\mathbf{N}$ points in the direction of radial unit vector $\mathbf{e}_r$, so, we write:

    \begin{equation}
    \mathbf{N} = |\mathbf{N}|\mathbf{e}_r.
    \end{equation}

    Friction $\mathbf{F}$ is in the opposite direction of the tangential unit vector $\mathbf{e}_\theta$, so, we need to place a minus-sign in its expression. Furthermore, as (dry) friction is usually modelled by the product of the coefficient of friction and the magnitude of the normal force, we can write:

    \begin{equation}
    \mathbf{F} = \mu|\mathbf{N}|(-\mathbf{e}_\theta).
    \end{equation}

    Lastly, weight is the force due to gravity, $|\mathbf{W}|=mg$, where $g$ is the gravitational constant. However, we need to express this force in terms of its components. In this case, those components are directed parallel to the radial and tangential unit vectors. As the latter are pointed (partly) upwards, as opposed to the downwards-pointing weight, we already know that both its components carry a minus-sign, i.e. $(-\mathbf{e}_r)$ and $(-\mathbf{e}_\theta)$. What remains, is the correct expression for the magnitude of the weight in terms of its respective unit vectors.

    To clearly show how we get an expression for $\mathbf{W}$ in terms of its components along the directions of $\mathbf{e}_r$ and $\mathbf{e}_\theta$, have a look at Figure 3.

    Figure 3. Finding the components of $\mathbf{W}$

    What you see is just the weight vector $\mathbf{W}$ from our force diagram in Figure 2, including the radial and tangential unit vectors $\mathbf{e}_r$ and $\mathbf{e}_\theta$. For visual clarity, we subtended them on mass $m$. Also added are the two component vectors in the opposite direction of the unit vectors for which we need to find expressions.

    Let component vector $\mathbf{v}_r = a(-\mathbf{e}_r)$ and component vector $\mathbf{v}_\theta = b(-\mathbf{e}_\theta)$, where $a$ and $b$ are some magnitude value such that the vector sum of $\mathbf{v}_r$ and $\mathbf{v}_\theta$ equals $\mathbf{W}$. In other words,

    \begin{equation}
    \mathbf{W} = \mathbf{v}_r + \mathbf{v}_\theta = a(-\mathbf{e}_r) + b(-\mathbf{e}_\theta).
    \end{equation}

    To find the values of the magnitude of $a$ and $b$, we use the fact that the magnitude $|\mathbf{W}| = mg$. So, using high school trigonometry, we deduce that

    \begin{align}
    a &= mg\sin\theta, \\
    b &= mg\cos\theta.
    \end{align}

    Now, we can write $\mathbf{W}$ in terms of its components by substituting equations (7) and (8) into (6):

    \begin{equation}
    \mathbf{W} = mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta).
    \end{equation}

    And so, if we substitute equations (3), (4), (5), and (9) into equation (2), we get:

    \begin{align}
    m\ddot{\mathbf{r}} &= |\mathbf{P}|\mathbf{e}_\theta + |\mathbf{N}|\mathbf{e}_r + \mu|\mathbf{N}|(-\mathbf{e}_\theta)\nonumber \\
    &\hspace{2em}+ mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta).
    \end{align}

    Step 4. Express the Cartesian $\ddot{\mathbf{r}}$ in polar coordinates

    As we know that the expression for the second time derivative of non-uniform circular motion is

    \begin{equation}
    \ddot{\mathbf{r}} = -R\dot{\theta}^2\mathbf{e}_r + R\ddot{\theta}\mathbf{e}_\theta,
    \end{equation}

    where $R$ is the radius of the circular motion, i.e. the cylinder. We proceed to substitute this into equation (10).

    And so, we get

    \begin{align*}
    m(-R\dot{\theta}^2\mathbf{e}_r + R\ddot{\theta}\mathbf{e}_\theta) &= |\mathbf{P}|\mathbf{e}_\theta + |\mathbf{N}|\mathbf{e}_r + \mu|\mathbf{N}|(-\mathbf{e}_\theta) \\
    &\hspace{2em}+ mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta),
    \end{align*}

    which, of course, after expansion, becomes

    \begin{align}
    -mR\dot{\theta}^2\mathbf{e}_r + mR\ddot{\theta}\mathbf{e}_\theta &= |\mathbf{P}|\mathbf{e}_\theta + |\mathbf{N}|\mathbf{e}_r + \mu|\mathbf{N}|(-\mathbf{e}_\theta) \nonumber \\
    &\hspace{2em}+ mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta).
    \end{align}

    Step 5. Resolve radially and tangentially

    We can now resolve equation (12) into its radial and tangential components.

    \begin{align}
    \mathbf{e}_r &: -mR\dot{\theta}^2 = N – mg\sin\theta, \\
    \mathbf{e}_\theta &: mR\ddot{\theta} = P – \mu N – mg\cos\theta.
    \end{align}

    Step 6. Write down the equation of motion (in polar coordinates)

    Rearranging equation (14), we can write down the second-order differential equation of motion:

    \begin{equation}
    \ddot{\theta} = \frac{P – \mu N – mg\cos\theta}{mR}.
    \end{equation}

    While we could have solved equation (14) for $N$, this would still leave us with the second time-derivative of $\theta$. Instead, we want an expression of $N$ in terms of a simple $\theta$. This means that we need to get rid of $\ddot{\theta}$ in some way. It is not immediately clear how equation (14) or (15) should be operated on to achieve this. However, here is a neat trick.

    Step 7. The trick

    Have a look at the following equation where we apply the chain rule:

    \begin{equation}
    \frac{\text{d}\dot{\theta}^2}{\text{d}t} = \frac{\text{d}\dot{\theta}^2}{\text{d}\dot{\theta}}\frac{\text{d}\dot{\theta}}{\text{d}t} = 2\dot{\theta}\frac{\text{d}\dot{\theta}}{\text{d}t} = 2\dot{\theta}\ddot{\theta}.
    \end{equation}

    So, if we substitute equation (15) into (16), we get

    \begin{equation}
    \frac{\text{d}\dot{\theta}^2}{\text{d}t} = 2\dot{\theta}\left(\frac{P – \mu N – mg\cos\theta}{mR}\right).
    \end{equation}

    If we now integrate both sides with respect to time, we get

    \begin{align}
    \int \frac{\text{d}\dot{\theta}^2}{\text{d}t}\text{d}t &= \int 2\dot{\theta}\left(\frac{P – \mu N – mg\cos\theta}{mR}\right)\text{d}t, \nonumber \\
    \dot{\theta}^2 + A &= 2 \int \frac{\text{d}\theta}{\text{d}t}\left(\frac{P – \mu N – mg\cos\theta}{mR}\right)\text{d}t, \nonumber \\
    &\text{where $A$ is an arbitrary constant}, \nonumber \\
    \dot{\theta}^2 + A &= 2 \int \left(\frac{P – \mu N – mg\cos\theta}{mR}\right)\text{d}\theta, \nonumber \\
    \dot{\theta}^2 + A &= \frac{2}{mR} \int (P – \mu N – mg\cos\theta)\,\text{d}\theta, \nonumber \\
    \dot{\theta}^2 + A &= \frac{2}{mR} \left( P\int 1\,\text{d}\theta – \mu N\int 1\,\text{d}\theta – mg\int \cos\theta\,\text{d}\theta\right), \nonumber \\
    \dot{\theta}^2 + A &= \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2mg\sin\theta}{mR} + B, \nonumber \\
    &\text{where $B$ is an arbitrary constant}, \nonumber \\
    \dot{\theta}^2 &= \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R} + B – A, \nonumber \\
    \dot{\theta}^2 &= \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R} + C, \\
    &\text{where $C=B-A$} \nonumber.
    \end{align}

    Solving the initial condition problem to find $C$, we use the fact that at $t=0$, angle $\theta = 0$, thus $\dot{\theta} = \ddot{\theta} = 0$. This renders $C = 0$ in equation (18), and so, we have

    \begin{equation}
    \dot{\theta}^2 = \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R}.
    \end{equation}

    Note, we now have obtained an expression for $\dot{\theta}^2$ which already appeared in equation (13). We can, therefore, substitute equation (19) in (13), and we obtain:

    \begin{equation}
    -mR\left(\frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R}\right) = N – mg\sin\theta.
    \end{equation}

    Expanding and rearranging this, we get

    \begin{align}
    N – mg\sin\theta &= -2P\theta + 2\mu N\theta + 2mg\sin\theta, \nonumber \\
    N – 2\mu N\theta &= -2P\theta + 2mg\sin\theta + mg\sin\theta, \nonumber \\ N(1 – 2\mu \theta) &= -2P\theta + 3mg\sin\theta, \nonumber \\
    N &= \frac{3mg\sin\theta – 2P\theta}{1-2\mu\theta}.
    \end{align}

    So, now we have an expression of $N$ in terms of the gravitational constant $g$, the variables $m$, $\mu$, and $P$, and the more reasonable $\theta$ instead of $\dot\theta^2$.

    And so, if we want to calculate when a mass would be slung out of its orbital path, we write $N = 0$ as this means, in physical terms, that the mass isn’t resting on the cylinder anymore (since it doesn’t exert a normal force on the mass). In other words, find the roots of equation (21) to find the one unknown variable. Note, $R$ does not play a role. Of course, bear in mind that $m$ is a point mass.

  • Why do wet clothes dry?

    Why do wet clothes dry?


    In the Northern Hemisphere, summer has arrived. The time has come for us to chase the general public with water guns, jump through the neighbour’s garden sprinklers’ water rays, and either carefully place a soggy, wet sea cucumber on a human’s belly during their beach nap(beginfootnote)The author does not approve of this. Sea cucumbers should be left alone.(endfootnote), or simply dump them (the human) in the actually-still-too-cold seawater, especially if you love them. At the end of the day, after all those wet adventures, nothing will beat hanging your clothes out to dry in a soothing breeze of fresh alpine air.

    A few years ago, a friend asked what exactly causes wet clothes to dry. How does that work, exactly? I thought it was a great question because what may seem like a simple problem actually exposes one of the fundamental aspects of the way our universe works and, at the same time, forms one of the main causes for headaches among undergraduates: the second law of thermodynamics.

    TL;DR? Don’t like mathematics (high school level) and prefer to read the ‘dashboard’ version? Skip to the bottom. Everyone else, please read on!


    A box of gas

    Let’s first paint ourselves a simpler picture than the actual situation where we wear whatever is the latest summer catwalk beach fashion. For now, we will also ignore the Sun, we will ignore the wind, and we will ignore the humidity of the air.

    Imagine, your colourful pair of swimming trunks is actually a simple box with a hundred gas molecules. The particles bounce chaotically back and forth against each other and the walls of the box itself.

    Now, let there be a hole in the wall. Imagine, by pure chance, one molecule escaping the box through the hole, arriving in another container of exactly the same size. This obviously means there are now only 99 gas molecules left in the original box.

    Figure 1. Two boxes with gas molecules bouncing around inside. In box (a), one has escaped, 99 remain. In box (b), 98 remain, two have escaped.
    Figure 1. Two pairs of boxes with gas molecules bouncing around inside. In (a), one molecule has escaped to the right box through the hole, 99 remain in the left box. In (b), two have escaped to the right box, 98 remain in the left box.

    Have a look at Figure 1a. Assuming all gas molecules look exactly alike, how many ways do we have to arrange them in order to get the same result? Well, instead of this particular molecule having escaped, any other one of these hundred molecules could have escaped just as well. And so, as each one of the hundred gas particles was capable of escaping the box, exactly a hundred possibilities could have led to the same outcome (i.e. 1 escaped, 99 remain). In other words, exactly one hundred different configurations, or microstates, will entail the microstate of the box where it lost one molecule while 99 remain inside. Let’s call this number $W$. And let’s call that number for the microstate where one molecule escaped (and 99 remain inside), $W(1)$. So, $W(1)=100$.

    Now, imagine not one, but two particles flew out, as is depicted in Figure 1b. Well, this means that a different number of arrangements would have led to this situation or microstate. As concluded above, for the first particle, one hundred possibilities existed as there were as many particles in the box, originally. For the second particle, however, only 99 possibilities existed since one had left the building already! Since for every 100 possibilities for the first molecule, 99 other possibilities exist for the second molecule, we calculate that the total number of possibilities leading to this particular state (i.e. 2 escaped, 98 remain), is $100 \times 99 = 9900$. However, since it doesn’t matter which of the two particles leaves the box first and which second, as they look exactly alike, we can divide that number by two, giving a total number of $4950$ possibilities. And so, $W(2) = 4950$.

    More accurately, in general, to calculate the possible combinations in a situation like this, we use the formula

    \begin{equation} W(k) = \frac{n!}{k!(n – k)!}, \end{equation}

    where $n$ is the number of molecules in the box initially, which is 100, and $k$ is the number of molecules having escaped through the hole. $W$ is the letter we will further use to denote the number of possible arrangements of our gas molecules for each situation (e.g. 0 escaped & 100 remain, $W(0)$, or 4 escaped & 96 remain, $W(4)$, et cetera).

    Here is a table with a few results. We included the situation where no single molecule has left the box. Obviously, the number of possible arrangements of the molecules leading to this situation, i.e. 0 escaped & 100 remain, is exactly one. We also included two more configurations where three and four particles have left the box. Note how quickly the possible arrangements increase.

    [table id=1 /]

    Probabilities

    How far can we take this? In this simplistic model, we can imagine the number of molecules remaining in the box becoming equal to the number having escaped into the other box: 50 escaped, 50 remain. So, let’s add $W(50)$ to the table. Also, let’s add more configurations to the table and have even more molecules escape the box until none are left, just to see what happens to the number of possible arrangements.

    [table id=2 /]

    As you can see, the number of possible arrangements decreases again after the box has reached its natural state of equilibrium (i.e. 50 escaped, 50 remain). This is, of course, only logical as the situation ‘flips’, as it were. More and more molecules end up escaping the box rather than remaining.

    If we were to calculate the probability of one or the other situations occurring, how should we go about this? Well, for instance, take the situation, or state, in which precisely zero gas molecules escaped. What is the probability of this state occurring?

    We would need to know the number of possible arrangements ($W$) in the state where there are 0 which escaped and 100 remain (that number is exactly 1, so, $W(0) = 1$), divided by the total of all possible arrangements in all states (the sum of all $W$’s). Mathematically, what we are calculating is the possibility $P$ where 0 molecules have escaped, in other words, for $P(0)$, we write:

    \begin{equation} P(0) = \frac{W(0)}{\text{total of all }W} = \frac{1}{\text{total of all }W}. \end{equation}

    Of course, the total of all $W$ still needs calculation. To do that, we use equation (1) to write down an equation for the sum of all $W$ in the previous table:

    \begin{align} \text{total of all }W &= \text{the sum of }W(0)\text{ to }W(100), \\ &= \sum_{n=0}^{100} \frac{100!}{n!(100-n)!}. \end{align}

    The answer to equation (4) is 1 267 650 600 228 229 401 496 703 205 376.

    That is a large number of total possible arrangements of all possible states. So, you can imagine that the probability of $P(0)$ occurring is inconceivably small: following equation (2), we get about $8 \times 10^{-29}$%. This would be 0% rounded to the nearest whole percentage.

    Likewise, we can calculate the probability of 50 escaping, 50 remaining, or P(50). This turns out to have a probability of 8% (to the nearest whole percentage). In the following plot of probabilities for each state, you can see which state is most likely to occur.

    A diagram showing that the equilibrium state (50 escaped, 50 remain) has the highest probability of just shy of 8%. Any other state has a drastically lower probability of occurring.

    So, in a way, given enough time, the arrangements of water molecules will converge to the state with the highest probability. Here, this is the situation where 50 molecules escaped, 50 remain. This is the so-called equilibrium point. Though there may be fluctuation around its equilibrium—plus or minus one or a few particles—it is perfectly fair to say that the probability of no molecules remaining, P(0), and the probability of all molecules escaping, P(100), are near zero. Intuitively, this is what you would expect: while it is theoretically possible, in practice, you will never live long enough to ever witness all the molecules randomly gathering in just one box.

    Back to our wet clothes

    In real life, however, our pair of swimming trunks are not a box nor are there only one hundred water molecules. There are billions of water molecules in a liquid phase held together by the fabric of our garment. Also, there is the Sun. And there might be wind or even just a slight breeze. Besides, not looking like a box, swimmers don’t have a hole attached to a second box.

    However, if the box is a metaphor for our wet clothes, then the second box is a metaphor for its environment, the open air. And now, it gets interesting.

    In our example of the two boxes, no further external forces played any role(beginfootnote)A system that is thermally isolated from its surroundings is called an adiabatic system.(endfootnote). There was no wind, no Sun, no air humidity to consider. In reality, of course, they should be taken into account. Here’s what they do: Sun heats up the water in the trunks, causing its molecules to gain energy, aiding escape from the fabric. Wind causes the air molecules to bump into water molecules, removing excess water vapour around the clothes, aiding the water molecules to evaporate even further. As long as the relative air humidity isn’t too high, drier air helps to evaporate the water even further.

    So, what do these circumstances do, exactly? They shift the equilibrium point in our plot to the right. States where more and more molecules escape and less remain in the box, that is, stay in the swimming trunks, get a higher probability of occurring due to these circumstances.

    In other words, if our boxes would be subjected to the elements, it would cause the equilibrium point to shift from 50 escaped, 50 remain, towards 95 escaped, 5 remain, for instance.

    Ludwig Boltzmann (1844-1906)
    Ludwig Boltzmann (1844-1906)

    Moreover, since the air outside is practically infinitely large compared to our swimming trunks—and not at all like the spatially limited second box of our metaphor—the number of ways water molecules can be arranged by random motion, in a system of clothes hanging in the air outside, is significantly leaning towards the state where most escape into the air, even without wind and sunshine. Even though it may take longer, that state is practically inevitable.

    It was the Austrian physicist Ludwig Boltzmann who elaborated on this very statistical nature of states of a system, specifically in terms of its possible configurations of microscopically small molecules per one and the same end state.

    Entropy and the second law of thermodynamics

    Grave of Ludwig Boltzmann on Zentralfriedhof (Central Cemetery), Vienna, Austria

    Now, because the number of possible arrangements, $W(k)$, gets very big very fast, Boltzmann calculated their natural logarithm value. This is a very neat function when dealing with incredibly large numbers and exponential growth. On any respectable high school calculator, this can be done using the ‘ln’-button.

    Boltzmann then proceeded to multiply these log values with a constant $k$(beginfootnote)which is a different $k$ than the one we used earlier. The value of this $k = 1.380649 \times 10^{-23}\text{JK}^{-1}$.(endfootnote) to link the phenomenon of mechanically mixing stuff (arrangements of molecules) with the thermodynamical phenomenon of entropy of heat. We now call this constant $k$ the Boltzmann constant. He then wrote down the famous expression which we now call Boltzmann’s equation for entropy. In Vienna, in the city’s Central Cemetery, his gravestone is engraved with this very formula:

    \begin{equation} S = k\log_e W. \end{equation}

    So, now we get a new table with values for entropy $S$:

    [table id=3 /]

    As you can see, entropy $S$ increases towards the equilibrium point, only to decrease beyond that, up to the point where it is zero again. Note that the highest value of entropy also has the highest probability value. This means that the state of the box and its surroundings (the other box) will tend to maximum entropy. This also means that an equilibrium point entails maximum entropy.

    Going back to our system of wet clothing and their surroundings (the air) this means that, here too, the state of wet clothing will tend to maximum entropy. This value will correspond to the situation where most of the water molecules have escaped the clothes.

    This is the second law of thermodynamics: the entropy of the Universe tends to a maximum.

    Why do wet clothes dry?

    While external factors such as sunshine, wind, and relatively low air humidity do cause the probability distribution to shift more towards the state where most water molecules escape the clothes, based on the random motion of molecules alone, statistically, they should leave the fabric anyway (even though this is a slower process than with sunshine, wind, et cetera).

    This is because the number of ways in which water molecules remain inside the clothes is simply almost infinitely small compared to the number of ways where they are not inside the clothes. This leads to the statistical fact that the probability of water molecules not remaining in the clothing outweighs the probability of them remaining in the clothing.

    There are simply more places for water molecules to be in the open air than there are within the constrained spatial dimensions of someone’s tight swimming trunks. Or anyone’s, really.

    Ultimately, wet clothes dry because the entropy of the Universe tends to a maximum.


    Photo of Boltzmann’s grave by Daderot, under CC BY-SA 3.0.

  • The formula that got Albert Einstein the Nobel Prize and should stop us getting sunburn all the time

    The formula that got Albert Einstein the Nobel Prize and should stop us getting sunburn all the time


    A copy of page 5 of the newspaper The Times of 10 November 1922. Near the bottom, a small article is printed. The title is Nobel Prize for Einstein. The text goes as follows. Stockholm, Nov 9.—The Nobel Prize for Physics—1921—has been awarded to Professor Albert Einstein, of Berlin, in recognition of his work in theoretical physics. The 1922 prize for physics has been awarded to Professor Niels Bohr, of Copenhagen, in recognition of his research work into the structure of atoms.—Reuter.
    ‘Nobel Prize for Einstein’, one sentence was spent in The Times of 10 November 1922.

    In 1921, Albert Einstein won the Nobel Prize “for his services to Theoretical Physics, and especially for his discovery of the law of the photoelectric effect.” Not a word about relativity. So, no, he did not win the Prize with $ E=mc^2 $. Though it is his most famous equation—which, by the way, is not the complete version—it is not his Nobel Prize-winning formula. We will write it down, but first, we describe what this photoelectric effect is.


    Different stuff is made up of different molecules. Different molecules are made up of different atoms. Different atoms are made up of a variety of nuclear composites and different numbers of electrons. So far, nothing new, perhaps, but here’s the thing. If electrons are exposed to particular amounts of energy, they can be ejected away from the nucleus.

    An atom of which one or more electrons have been blasted away is called an ion. The process is called ionisation. Whether this occurs, depends on a few things such as the type of stuff (=the type of atoms and how they are bound together) and the specific energy it is exposed to.

    If ionisation at the surface of a material is achieved by normal light, we call this the photoelectric effect: light (the ‘photo’-part) causing electrons to leave their nucleus (the ‘electric’-part).

    A diagram of the ionisation of an atom (not to scale). (1) The yellow cloud represents an electron’s (probable) whereabouts. The tiny pink core represents the atom’s nucleus. (2) Photons of a specific colour radiate towards the atom. (3) The electron has flown off. The nucleus remains. The atom has become an ion.

    Not about intensity

    One peculiar thing is worth mentioning. In fact, it was this puzzle that led Albert to his equation. It turned out that what matters is the frequency of the light beam, i.e. the colour of the light, not the intensity of it, i.e. the power per square metre, or Joule per second (watt) per square metre.

    Imagine, in the diagram above, that a billion yellow photons would radiate towards the atom and nothing happened; the electron would stay where it was. Now imagine a billion billion billion billion yellow photons approaching the atom. Still nothing would happen as it is not about intensity.

    Yellow light is less energetic than blue light, so if you would replace the light bulb for a source that delivers pure blue light, with even one blue photon, it could happen easily (though you would have to aim impossibly precise, so it makes sense to actually radiate a lot). This puzzled many scientists, but Albert solved it and won the Nobel Prize.

    With his discovery, quantum physics was starting to get momentum. He, and other good physicists of his time, showed that light could be seen as little packets of energy, which scientists started calling photons. A beam of light was now a stream of photons. The intensity, the amount of photons per second per square metres doesn’t matter but the frequency of a photon, or energy per photon does.

    DNA

    While this is all cool and useful for scientific purposes, we certainly do not want any electrons of the DNA molecules of our skin breaking away from their atomic confines. Atomic bonds would be destroyed and our DNA would become mutated. Even though astonishing molecular biological processes in our body repair defects like this in a staggering, basically inconceivable number of cases, some errors might slip through and may even become the start of tumour growth. Therefore, it is important to know what energy domains would cause our beloved bodily electrons to be blasted off so that humanity can learn to avoid those dangerous environments.

    The problem arises when we get into the mid to high-energy electromagnetic radiation, or light, or photons, if you will. We’re talking the dangerous kind of ultraviolet here, the type of UV causing DNA mutation to occur: UVB to be precise. A photon of UVB-light is about 1.8 times more energetic than a photon of the yellowish light in your home and almost a million times more energetic than a mobile phone photon. So, don’t be scared of being home. As soon as you set foot outside, though, be afraid. Not of the dark, but of the light, for ionising UVB-light is emitted by the Sun.

    A diagram of electromagnetic radiation. Far right, we see the dangerous types of radiation: cosmic rays, x-rays, gamma rays, UV-light. In the middle, we see visible light. Far left, we see the lowest energy photons: WiFi, mobile phones, microwave ovens.
    A diagram (not to scale) of electromagnetic radiation, or photons, if you will. The mentioned values are the frequencies of the photons, expressed in gigahertz (GHz). The higher the frequency, the higher the energy of the photon.

    Fortunately, as stated before, our bodies have evolved to repair the damage when necessary. This is why even X-rays are okay and hospitals and dentists make sure not to expose you to doses of energetic photons you wouldn’t survive. Continuous monitoring of its uses and effects is prerequisite.

    It’s partly a question of the law of large numbers, though. If the number of freely whizzing electrons is large enough, they themselves will become the main cause of an increasing number of damaged DNA molecules, and, eventually, some repairs will fail or not even take place. So, while it is not instantly dangerous, we do recommend some reading up on the subject of sunbathing. Use UV protection. Don’t get sunburnt. And give your body a chance to recover from the ruthless blasts of ionising UV radiation. Forget microwaves, the problem is crispy skin.

    The formula

    So, now we finally get to Albert’s Nobel Prize-winning formula. Here it is

    \[ \frac{1}{2}m_ev^2_\text{max} = h\nu – \phi. \]

    It doesn’t look as sassy as the other one, right? And yet, it’s the one that allows us to calculate if, for instance, electrons of our body’s carbon atoms get blasted out by the photons emitted by the lamp in your lavatory (they do not). Or if the laser pointer knocks some electrons out (it doesn’t), which we use anyway, because we need to point at things on our PowerPoint slides as they might well be ill-designed (they are).

    So, $ \frac{1}{2}m_ev^2_\text{max} $ means maximum kinetic energy, which is simply the energy with which an electron flies away from its nucleus. If its value turns out to be smaller than or equal to zero then the electron is not affected at all. It’ll keep stuck to its nucleus. If it is larger than zero then off it goes. The symbol $ h $ is a constant, which we needn’t worry too much about. It’s a number and it’s called the Planck constant. The Greek letter $ \nu $ is the frequency of the photon. In the diagram above, a few have been mentioned. Mind you, $ h\nu $ means $ h \times \nu $ and is the energy of a photon. Mathematicians, physicists, engineers, and other folks, just like to leave out the $ \times $-sign. The Greek letter $ \phi $ is the so-called work function. It is the minimal energy needed for the occurrence of a photoelectric effect. Its value depends on the type of atom, molecule, material, and surface you want to calculate the photoelectric effect of.

    In conclusion

    Notice Einstein’s formula does not have any term relating to the number of photons radiated per second per square metre towards the atom of interest, i.e. the intensity. Only the frequency is important. This means that atoms—such as your body—will be left undisturbed irrespective of the power of the radiation they are exposed to. There may be a bit of heat but there is no ionisation. The potential danger lies in frequency ($ \nu $), such as that of UV light and higher. Here, both dosage and capability of recovery play a crucial role.

    Young Albert Einstein

    The value of the Planck constant is $ h = 6.626070 \times 10^{-34} $ Js (Joulesecond). The value of the work function of carbon, of which our entire body is made, including our DNA, is $ \phi = 8.0108831 \times 10^{-19} $ J. If a WiFi photon has a frequency of 2.5 GHz, you can calculate yourself if it would yank the electrons from a carbon atom. Remember to convert 2.5 GHz to $ 2.5 \times 10^9 $ / s (per second). Thanks to Albert, calculating this has become child’s play. We could do the maths on the back of an envelope. If all the terms on the right hand side of the equal sign turn out to be larger than zero, then sell your router immediately and—based on this diagram—you most definitely ought to refrain from switching on the light while frequenting the lavatory. Good luck with the calculation! (Or check the working out.)


    Featured image: a 14-year-old Albert Einstein, photographed in 1893. Credits EMILIO SEGRE VISUAL ARCHIVES / AMERICAN INSTITUTE OF PHYSICS / SCIENCE PHOTO LIBRARY / Universal Images Group. Source: Young Albert Einstein, physicist. [Photography]. Encyclopædia Britannica ImageQuest. Retrieved 9 Mar 2019, from 
    https://quest.eb.com/search/132_1258083/1/132_1258083/cite

    Smaller image of an even younger Albert Einstein: Credits EMILIO SEGRE VISUAL ARCHIVES / AMERICAN INSTITUTE OF PHYSICS / SCIENCE PHOTO LIBRARY / Universal Images Group. Source: Young Albert Einstein, physicist. [Photography]. Encyclopædia Britannica ImageQuest. Retrieved 9 Mar 2019, from https://quest.eb.com/search/132_1255429/1/132_1255429/cite

    Newspaper article: “Nobel Prize for Einstein.” Times, 10 Nov. 1922, p. 5. The Times Digital Archive. Retrieved 8 Mar 2019 from http://tinyurl.galegroup.com/tinyurl/9Q37o0.

  • What is a spacetime interval?

    What is a spacetime interval?


    Einstein and collaborators taught us that space and time are not fixed quantities. They can stretch and contract. They vary. There is one thing, though, that does not vary. It is the invariance of the spacetime interval.

    Download PDF


    Spatial interval

    Suppose, a photon is emitted from origin $O$ and travels to point $F$ as depicted in Figure 1. Let us write down the expression for its distance-squared, $d(O,F)^2$, in terms of the other distances using the good old Pythagorean theorem:

    $ d(O,F)^2 = d(O,A)^2 + d(A,B)^2 + d(B,F)^2. $

    Figure 1 A photon travels from O to F in a three-dimensional space

    We can also write the previous expression in terms of their distance from $O$. We then write the following:

    \begin{align}
    F &= (x, y, z),\quad O = (0,0,0), \\
    d(O,F)^2 &= (x-0)^2 + (y-0)^2 + (z-0)^2, \\
    \therefore d(O,F)^2 &= (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2. \label{eq:distance O-F}
    \end{align}

    As the axes of the space in Figure 1 are spatial and Euclidean, $d(O,F)^2$ is called a spatial or Euclidean interval. It can also be thought of as a rectangular cuboid represented by its space diagonal $d(O,F)$, tracing out a region of 3D space.

    In the real world, to make sure we meet at the correct place, we could, for instance, give the following coordinates: 1 Einstein Drive, 2nd floor. Think of Einstein Drive as some place along the the $x$-axis (next to $x$-axis places like Battle Road, Mercer Road), number 1 as some place along the $y$-axis, and 2nd floor as some place along the $z$-axis.

    What we still need, though, is an extra bit of information: when do we meet?

    Time interval

    Suppose, Figure 2 shows a timed series of our photon on its way to point $F$ and beyond. It demonstrates that we live in a world where we do not just need three spatial coordinates, but also a time coordinate. It is only logical to not just tell the people you are supposed to meet, where in space you will be, but also when in time you will be there.

    Our photon $P$ flies through $F$ at $t=3$. This entails that the time coordinate of the event that the photon reaches $F$ is
    $ t_{F}=3. $

    Assuming that at $t=0$, photon $P$ is at the origin,

    $ t_{O}=0, $
    then we can write for the temporal interval between the photon leaving $O$ and reaching $F$:

    $ \Delta t_{OF} = t_{F} – t_{O} = 3 – 0 = 3. $

    Figure 2 A photon travels from O to F in a three-dimensional space over a period of time

    Time to distance unit conversion

    As the previous two sections showed, we need four coordinates to describe an event, for instance, the event where photon $P$ reaches $F$. The three spatial distances are measured in a unit of distance, usually, metres. The one temporal distance is not a distance in the traditional sense and is usually expressed in seconds. This makes it difficult to make sensible comparisons.

    To convert the time-units to distance-units, we multiply by a constant of nature, the speed of light $c$, which, by Einstein’s second postulate [1], happens to be invariant: no matter which frame of reference you choose, the speed of light is constant. For a longer description of this conversion, read section 4.3 of Deriving the Lorentz transformations from a rotation of frames of reference about their origin with real time Wick-rotated to imaginary time. We conclude that our time interval becomes a temporal distance:

    $ \Delta t \mapsto c\Delta t. $

    Spacetime interval

    In Figure 3, we left out the spatial $z$-axis and replaced it with the temporal $ct$-axis (which is thus time expressed in distance-units) in order to make a comprehensible drawing on a flat surface. In reality, of course, the photon still moves in the $z$-direction as well. (We have thus not yet been successful to draw a four-dimensional object on a flat surface.) Mind the unit vector diagram top right and the points of distances $\Delta x$, $\Delta y$, and $c\Delta t$. Then think, really hard, of an added fourth spatial distance $\Delta z$, somewhere.

    To calculate the spatial distance $d(O,F)$ for our photon, we look again at Equation \eqref{eq:distance O-F}:

    $ d(O,F)^2 = (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2.\quad\eqref{eq:distance O-F} $

    Since we know that speed, in general, is calculated through $v = \Delta x / \Delta t$, where $x$ is the travelled distance in one direction, and $v = c$ for our photon, we can write for the travelled distance of our photon from $O$ to $F$:

    \begin{align} d(O,F) &= v\Delta t, \\ d(O,F)^2 &= (v\Delta t)^2, \\ \therefore d(O,F)^2 &= (c\Delta t)^2. \end{align}

    This is becoming interesting, since $(c\Delta t)^2$ is also (the square of) the temporal distance. If we substitute Equation \eqref{eq:distance O-F} into this last equation, we get

    $ (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2 = (c\Delta t)^2. $

    If we rearrange this a little bit, we get

    \begin{equation}
    – (c\Delta t)^2 + (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2 = 0. \label{eq:spacetime-homogeneity}
    \end{equation}

    While this may seem nice and simple, the question we should be asking is, what is zero? If we know the answer to that, we know the answer to what all the terms are on the left-hand side of the equals sign.

    In physics and mathematics, whenever something equals zero, something special is going on: it may entail a certain system in a certain configuration that is stable, static even, it may point to constant motion, an energy well, an attractor, a root, a conservation law, homogeneity, or a minimum or maximum of some kind.

    In general, it means that there is a certain kind of symmetry at play, which in turn means that something is conserved. There are beautiful, deep insights to be made as Emmy Noether showed us [2], and her genius deserves nothing less than an entire series of articles on their own.

    However, for now, let us conclude that independent of which coordinate system one uses, rendering different values for $\Delta x$, $\Delta y$, $\Delta z$, and even $\Delta t$, as we have come to learn from the Lorentz transformations, the sum of all these variances remains invariant. The zero points to the fact that irrespective of its four moving parts—no matter what frame of reference you prefer—the resultant is a constant, i.e. invariant.

    The quantity on the left-hand side has a name; it is called the spacetime interval and is denoted by $(\Delta s)^2$. The $s$ stands for ‘separation’. It is about the separation between events. If we had used the word distance, it might have had inadvertently referred too much to a spatial distance, hence, we use separation, $(\Delta s)^2$. And so, the spacetime interval is usually written:

    \begin{equation}
    (\Delta s)^2 = – (c\Delta t)^2 + (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2.
    \end{equation}

    The signs before the terms may have been flipped in some texts, but important to note is that, while time has been made comparable to space unit-wise by multiplication by $c$, you can still see that time has a special place in the interval of the fabric of the cosmos.

    Figure 3 A spacetime diagram with two spatial dimensions and one temporal dimension.

    Featured image: Klaus P. Rausch

    References
    1. A. Einstein, Zur Elektrodynamik bewegter Körper, Annalen der Physik 322(1905), no. 10, 891—921.
    2. E. Noether, Invariante Variationsprobleme, Nachrichten von der Gesellschaft der Wissenschaften zu Göttingen, Mathematisch-Physikalische Klasse 1918(1918), 235—257.
  • Deriving the Lorentz transformations from a rotation of frames of reference about their origin with real time Wick-rotated to imaginary time

    Deriving the Lorentz transformations from a rotation of frames of reference about their origin with real time Wick-rotated to imaginary time


    Well-known for their central role in Einstein’s Special Relativity, the Lorentz transformations are derived from the rotation of two frames of reference in standard configuration while time is taken to be an imaginary unit of spacetime. This is rarely seen in the wild. Not many undergraduate textbooks or online texts show the details of the working. Hence, this article.

    Download PDF

    Introduction

    One might think this means that imaginary numbers are just a mathematical game having nothing to do with the real world. (…) It turns out that a mathematical model involving imaginary time predicts not only effects we have already observed but also effects we have not been able to measure yet nevertheless believe in for other reasons. So what is real and what is imaginary? Is the distinction just in our minds?

    S. Hawking[1]

    Even though there are many derivations of the Lorentz transformations to be found in textbooks, in syllabi, and online, to me, one of the most elegant remains the version Henri Poincaré once alluded to[2], which Hermann Minkowski then toyed with a bit further—to put it unreasonably mildly—in what we now call Minkowski space, but is rarely expounded in the aforementioned places.

    Henri Poincaré noted that, when the time axis of the two coordinate systems has been made imaginary, i.e. the imaginary axis in the complex plane, the transformations set forth by Hendrik Lorentz pop out automatically after a rotation of two reference frames in that complex plane.

    In this document, we show how this is done. We assume the reader is familiar with complex numbers.

    The aim is to derive the following set of Lorentz transformations:

    \begin{align}
    t’ &= \frac{t-vx/c^2}{\sqrt{1-v^2/c^2}}, \label{eq:Lorentz t-prime} \\
    x’ &= \frac{x-vt}{\sqrt{1-v^2/c^2}}, \label{eq:Lorentz x-prime} \\
    y’ &= y, \\
    z’ &= z,
    \end{align}

    where $(t,x,y,z)$ and $(t’,x’,y’,z’)$ are the coordinates of an event in two frames. The primed ($’$) frame is, seen from the unprimed frame, moving with speed $v$ in the $x$-direction. The speed of light in a vacuum is denoted by $c$. As a side note, the recurring term $(\sqrt{1-v^2/c^2})^{-1}$ is called the Lorentz factor and it is usually denoted by the letter $\gamma$.

    Standard configuration

    Figure 1: Two frames of reference in standard configuration. A two-dimensional depiction of Hermann Minkowski’s frame M and Albert Einstein’s frame E in standard configuration: the latter moves at speed v relative to the first in the direction of x. There is no motion in either the y- or z-direction. Note that some time has passed in this diagram. At time t=0, however, their origins were equal. In other words, at temporal coordinates t=t’=0, their spatial coordinates were the same, thus x=x’=0.

    Suppose, Hermann is standing still on the ground. Albert is driving his car and moves away from Hermann at speed $v$. We then have two frames of reference. There is Hermann’s frame $\mathcal{M}$ (the ground), with its origin $O$ at Hermann’s feet on the ground. And there is Albert’s frame $\mathcal{E}$ (the car), with its origin at Albert’s bottom on his chair. Their frames of reference are said to be in standard configuration as depicted by Figure 1. This means that at time $t=0$ in frame $\mathcal{M}$ where $x=0$ as well, the time $t’=0$ and position $x’=0$ in frame $\mathcal{E}$, too, and that one frame is in uniform (constant) motion relative to the other. In other words, $\mathcal{M}$ and $\mathcal{E}$ are said to be synchronised when the spacetime coordinates

    \[ (t,x) = (t’,x’) = (0,0). \]

    Of course, in the real world, there are four spacetime coordinates for each frame, i.e. $(t,x,y,z)$ and $(t’,x’,y’,z’)$, but to make our calculations a little bit easier, we consider the temporal coordinate $t$ (and $t’$) and spatial coordinate $x$ (and $x’$) only.

    So, looking at Figure 1, we can see from Hermann’s point of view – standing in the origin $O$ of $\mathcal{M}$ – that $\mathcal{E}$’s origin $O$ moves at speed $v$ relative to the $x$-axis of $\mathcal{M}$. Speed $v$, of course, just means that $\mathcal{E}$ moves at a certain amount of units of $x$ (say, metres) per a certain amount of units of $t$ (say, seconds). This is nothing new, but it is for our derivation of the Lorentz transformations important to repeat our secondary education for a little bit:

    \[ v = \frac{\Delta x}{\Delta t}, \]

    in Hermann’s frame of reference $\mathcal{M}$. More or less conversely, if we want to calculate how many spatial units frame $\mathcal{E}$’s origin has moved from frame $\mathcal{M}$’s origin, we rewrite the last equation into the perhaps more familiar law of uniform motion:

    \begin{equation}
    \Delta x = v \Delta t,
    \label{eq:x=vt}
    \end{equation}

    in Hermann’s frame of reference $\mathcal{M}$.

    As a side note, do realise that to Albert, his car is not moving at all; he is sitting in it. (Rather, it is the rest of the world that is moving with respect to his car and himself.) If the car were moving with respect to Albert, an accident with potentially serious consequences would be impending. So, for Albert’s sake, his speed within his own frame $\mathcal{E}$ (the car) is expressed as $v’=0$, as long as he stays put and buckled up in his chair. And so, the law of uniform motion of Albert, within his frame $\mathcal{E}$, becomes:

    \[ \Delta x’ = v’ \Delta t’ = 0 \Delta t’=0. \]

    Invariances

    If $\mathcal{E}$ is in constant motion with respect to $\mathcal{M}$, in one direction, the $x$-direction, as expressed in Equation \eqref{eq:x=vt}, then, mathematically, we call this a geometric translation in the $x$-direction. In physics, it is called a translational motion in the $x$-direction.

    Figure 2: Albert fires a photon. At time t=t’=0, Albert fires a photon P into direction x. Both the photon and Einstein’s frame E move into that same x-direction.

    Suppose, at time $t=t’=0$, Albert activates his special on-board photoelectric cannon, firing exactly one photon $P$ in the $x$-direction. Figure 2 shows how the photon is travelling through the spaces of both frames of reference.

    Looking at the diagram, we see that the spatial coordinates in the $y$-direction remain unchanged, $y=y’=0$, so we leave this out of our equations further on, to keep it simple. However, since $\mathcal{E}$ is moving with respect to $\mathcal{M}$ in the $x$-direction, we do know that $x\neq x’$ for $t>0$. And since we do not know for certain that $t=t’$ for $t>0$, only that $t=t’=0$, we will have to conclude that the position of $P$ differs:

    \begin{equation}
    \begin{aligned}
    \text{in Albert’s }\mathcal{E}\text{: }P &= (t’,x’), \\
    \text{in Hermann’s }\mathcal{M}\text{: }P &= (t,x).
    \end{aligned} \label{eq:coordinates of P}
    \end{equation}

    Fortunately, accepting Einstein’s Voraussetzungen[3], we know that the speed of light, $c$, is the same for every frame of reference. Using Equation \eqref{eq:x=vt}, $x=vt$, and the fact that $v=c$, in this case, we can write for the distance travelled of photon $P$ – the yellow line in the diagram – in the coordinates of the respective frames of reference:
    \begin{align}
    \text{in Albert’s }\mathcal{E}\text{: }\Delta x’ &= c\Delta t’, \\
    \text{in Hermann’s }\mathcal{M}\text{: }\Delta x &= c\Delta t.
    \end{align}

    As it is possible for any coordinate system to have points which lie on the negative side of the origin of a spatial dimension such as $x$ in our case, and thus for light to travel in the negative $x$-direction, we simply square both equations to always obtain a positive value.

    \begin{align}
    (\Delta x’)^2 &= (c\Delta t’)^2, \\
    (\Delta x)^2 &= (c\Delta t)^2.
    \end{align}

    If we then rearrange this,

    \begin{align}
    (\Delta x’)^2 – (c\Delta t’)^2 &= 0, \\
    (\Delta x)^2 – (c\Delta t)^2 &= 0,
    \end{align}

    we see that both are equal to zero, allowing us to write

    \begin{equation}
    (\Delta x’)^2 – (c\Delta t’)^2 = (\Delta x)^2 – (c\Delta t)^2.
    \label{eq:interval}
    \end{equation}

    This is a beautiful result, because it tells us that no matter what frame of reference you happen to be in, besides $c$, Albert and Hermann agree on the quantity $(\Delta x)^2 – (c\Delta t)^2$, despite the fact that the coordinates of $P$ are not necessarily the same in every frame of reference as we saw in \eqref{eq:coordinates of P}. In other words, both $c$ and $(\Delta x)^2 – (c\Delta t)^2$ are said to be invariant.

    You might wonder, what is this invariant quantity $(\Delta x)^2 – (c\Delta t)^2$, exactly? Well, this will be discussed in another post called What is a spacetime interval? And now, we might have just told you what it is. Anyway, let us move on to deriving the Lorentz transformations, and just keep in mind that $(\Delta x)^2 – (c\Delta t)^2$ is a wonderfully invariant quantity, equal in both frames of reference. Let us move on to imaginary time.

    Wick rotation and imaginary time

    Number sets

    Figure 3: real number line. A segment of the real number line, the set R of all real numbers, which goes on to infinity on either side.

    As many of us should know, Figure 3 depicts (a segment of) the real number line, that is the set $\mathbb{R}$ of all real numbers. It formed the culmination of all the previous extensions of the then-known set of numbers. Starting with the natural numbers, a set usually denoted by $\mathbb{N}$, containing all positive integers, arisen from the natural act of counting, the numeric repertoire was then extended by the notion of negative integers. Instead of just 1,2,3, we could now also count to -1,-2,-3 etc. This extension is denoted by $\mathbb{Z}$. Needless to say that $\mathbb{N}\subset\mathbb{Z}$, but we just did anyway.

    Of course, some people were clever, acknowledging the need for another extension: numbers which represented ratios, better known as rational numbers, such as 1/2, 1/-3, 1/4, -1/100, in other words, quotients of two integers. These numbers would sit in-between the integers in $\mathbb{Z}$. The symbol is $\mathbb{Q}$, and it is superfluous to add that $\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}$.

    While specified on a tablet, found in Susa (Iraq) in 1936, dated as used by Babylonians around 2000 BCE, that

    \[ \frac{3}{\pi}=\frac{57}{60}+\frac{36}{(60)^2}, \therefore \pi = \frac{25}{8}=3.125, \]

    it wasn’t until 1761 that a proof that $\pi$ is irrational was found by Johann Heinrich Lambert[3] meaning that it could not be constructed by any ratio of integers, as were many other numbers, such as $\sqrt{2}$. And so, yet again, an extension of the existing number line was needed. This was the aforementioned line representing the set $\mathbb{R}$, or, to be precise, $\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}$.

    And then, in the 16th century, people such as rivals Tartaglia and Cardano independently recognised that solutions to cubic equations sometimes required the manipulation of square roots of negative numbers, such as $\sqrt{-1}$. Later, Bombelli developed proper operations such as addition and subtraction. A whole slew of subsequent mathematicians then developed over several decennia what is now known as the complex plane or gaussian plane[5], representing the set $\mathbb{C}$, extending the real number line with an imaginary axis with multiples of the imaginary unit $i=\sqrt{-1}$. (It is, obviously, the solution to the quadratic $x^2+1=0$.) We realise mentioning that $\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}\subset\mathbb{C}$ is utterly redundant at this point.

    Translation and rotation

    Figure 4: Number sets. Every consecutive number set is an extension of the previous one. We can move from a simpler set to a more complex one, for instance, by simply multiplying our current position by a number only present in the more complex set. Although it seems like we are tumbling from one set to another, we are really just ‘sliding left or right’, one-dimensionally, on the number line of the more complex set. This sliding is called a translation. Note: the amount of ticks in Q is much larger, but for obvious reasons of legibility, we only ticked every 1/2-ratio.

    Let us have another look at the natural number line of $\mathbb{N}$. If we would want to convert the number 1 to a number that could not exist in $\mathbb{N}$ but could exist on the integer number line of $\mathbb{Z}$, let us then simply multiply the natural number 1 with a number from $\mathbb{Z}$, the negative integer $-1$. Since $1\times-1=-1$, we have transitioned from $\mathbb{N}$ to $\mathbb{Z}$. We ‘slid’ from 1 to $-1$, albeit in a different number set, which, mathematically, is the same as a translation by $-2$. This is easily expressed as starting from position 1, adding $-2$, and ending up at position $-1$ on the number line of, at least, $\mathbb{Z}$ (but not $\mathbb{N}$): $1+-2=-1$. Figure 4a aims to depict this.

    We can do the same with moving from position $-1$ in $\mathbb{Z}$ to a number not in $\mathbb{N}$, nor in $\mathbb{Z}$, but at least in $\mathbb{Q}$ by simply multiplying by a fraction, such as $-1/2$, which is also a number not in $\mathbb{N}$, nor in $\mathbb{Z}$. This is, again, actually a translation, though now by adding $3/2$: $-1+3/2=1/2$. Figure 4b aims to depict this.

    Similarly, transforming from position 1 in $\mathbb{Q}$ to $\mathbb{R}$, we multiply by, for instance, $\sqrt{2}$, which exists in $\mathbb{R}$ but not in $\mathbb{Q}$, and so, the result, $1\times\sqrt{2}=\sqrt{2}$ is in at least $\mathbb{R}$ but not in $\mathbb{Q}$, nor in $\mathbb{Z}$, nor in $\mathbb{N}$. The result is also a translation of $1+(\sqrt{2}-1)=\sqrt(2)$ in $\mathbb{R}$. Figure 4c aims to depict this.

    Figure 5: Wick rotation of 1 and real time. Compared to the increasing complexity of the number lines in Figure 4, this one is the most complex so far. Two Wick rotations into the complex (number) plane C, where (a) the Re-axis is the real number line of the set R and the Im-axis is the imaginary unit line of the set C. Note that a complex number consists of both: a real part and an imaginary part. For instance, the complex number z is written in the form z = a + bi, with i = √-1. The real part of z is Re(z) = a, and the imaginary part Im(z) = b. And so, z = 1 + i, z = 1/2 + 3i, z = 3, are all complex numbers, where the latter has an imaginary part of Im(z) = 0, which we simply leave out as 0i = 0. A complex number is thus two-dimensional, embedded in a plane with a real axis and an imaginary axis. (b) Wick-rotating all real numbers on the time axis to the imaginary axis, transforms real time into imaginary time.

    Note that, so far, the transformation of 1 or another number has involved a simple ‘sliding’ motion on the number lines, that is, one-dimensionally. Every time a new kind of number was introduced – the negative integers, ratios, and, lastly, the real numbers – a new set of numbers was created, and the number line evolved from discrete ($\mathbb{N}$) to a line continuum $\mathbb{R}$. The question is, what would be the next extension and what would it look like?

    As stated earlier, in the 16th century it became clear that a new type of number was necessary to solve a slew of quadratic equations. Owing to people such as Wallis, Wessel, Argand, Buée, Mourey, Warren, Français, Bellavitis, Gauss, and Euler[5][6], the idea to extend the real number line of $\mathbb{R}$ with an imaginary, perpendicular number line came to fruition. This created the so-called complex (geometric) plane, sometimes called the $z$-plane, Gauss plane or Argand plane. It is important to note that transforming a real number in $\mathbb{R}$ to a complex number in $\mathbb{C}$ involves not a translation but a rotation. Multiplying a real number in $\mathbb{R}$, say 1, by a number that only exists in $\mathbb{C}$, say $i$, is the same as geometrically rotating our position 1 on the real axes by $\pi/2$ onto a position $i$ on the imaginary axis as is depicted in Figure 5a.

    What if we did this with the entire real time axis in a space-time diagram as shown in Figure 5b? Every element of real time $t$ is multiplied by $i$. In other words, every part is rotated in the complex plane to become an entire imaginary axis of time $it$. This procedure is called a Wick rotation, named after theoretical physicist Gian Carlo Wick, who described such a procedure to solve problems in quantum and statistical mechanics[7].

    This seems promising and is what Henri Poincaré alluded to fifty years earlier. Before we continue, we have to do one other little thing. It has something to do with units of measurement.

    Minkowski diagrams

    Figure 6: The distance-time-diagram we all grew up with. The independent variable time t as the x-axis, and the dependent variable distance, x, as the y-axis. Four particles are travelling through time and (one-dimensional) space, each with its own distance function of time, that is, each with its own speed. Note that P travels the most amount of distance over the same period, in other words, it is the fastest. Note that S travels exactly zero distance in that same amount of time.

    We all grew up learning to read and use a type of diagram as depicted in Figure 6 during our physics classes. Time is put on the $x$-axis and distance $x$ is put on the $y$-axis. Somewhat confusingly, at first, as one might have gotten accustomed to using values of $x$ on the $x$-axis during the maths lessons. Of course, one learns that it is less about the names of variables and axes, rather, it is a matter of which is the independent and which is the dependent variable. The independent one, in this case, time $t$ (time flies, whether we want to or not), is then laid out over the axis called $x$ (which has not much to do with the variable named $x$), and the dependent one, a variable which happened to be named $x$, is projected onto the $y$-axis.

    In this diagram, we see four particles. The fastest, $P$, is moving away in the $x$-direction (which is up, but not necessarily up into the sky, do realise that!) covering more units of $x$ than the other three after the same time $t_1$ has passed. This is why it has a steeper slope. The slowest one is the one that is not moving at all, the stationary particle $S$. It is moving in time, which is why it exists at time $t_1$, but, spatially, it does not exist at a certain amount of units of $x$ away from the origin. In fact, it exists in exactly the same place, the origin.

    Figure 7: tx-diagram. A physicist’s diagram, where distance in space, x, is projected on the x-axis and distance in time t is projected on the y-axis. Note that the faster a particle travels, the smaller the angle of its ‘line’ through space and time with the x-axis. If the particle is stationary, it only ‘travels’ through time and the angle with the x-axis is maximised at π/2. In other words, it just goes straight up.

    Well, get yourself out of the habit: turns out that professional physicists like to flip the axes when it comes to time. In other words, they project the distance variable $x$ onto the $x$-axis, while the time variable $t$ almost invariably gets to be projected onto the $y$-axis. Yes, you heard it correctly. Time goes up in a physicist’s diagram. The esteemed professor Leonard Susskind, a theoretical physicist at Stanford University, even postulated, in part jokingly, during a lecture on the principle of least action that physicists are the only type of people who do this(beginfootnote)See, for instance, https://youtu.be/3apIZCpmdls?t=1447(endfootnote). And so, we flipped our diagram as you can see in Figure 7.

    Speaking of units, usually, time is measured in seconds and distance in metres. Usually. Though, remember when you went to visit those new friends of your parents and that one of the first things they assured their hosts is that their hometown was actually not that distant and that it was just ‘a two-hour drive’? Distance, while usually measured in kilometres between two places, is now expressed in units of time. Assuming that people legally drive – from door to door – at an average speed of $100\text{ km/h}$, the distance will be around 200 km.

    Why do people like to express distance in terms of units of time sometimes? Well, in some cases, people aren’t interested in the exact amount of kilometres, but rather tend to focus on how much of our valuable time a certain activity consumes, hence, an answer in units of time makes sense.

    Astrophysicists do another interesting distance-as-time conversion when it comes to distances between galaxies, for instance. They work with visible light reaching their telescopes, and other types of radiation. Moreover, the distances they work with are ridiculously large, especially when expressed in kilometres. So, they work with light-years, which sounds like a unit of time, but denotes a certain distance. One light-year is the distance light travels in one Julian year, which is $365.25$ days. Light travels at a speed of $c=299792458\text{ ms}^{-1}$ in the vacuum. To calculate the number of seconds in a Julian year, we multiply the number of seconds in one minute times the number of minutes in one hour times the number of hours in one day times the number of days in one Julian year:

    \begin{align}
    60\text{ s} &\times 60\text{ minutes} \times 24\text{ hours} \times 365.25\text{ days} \\
    &= 31557600\text{ s}.
    \end{align}

    Using Equation \eqref{eq:x=vt} to calculate distance $x$ light travels in one Julian year, we get

    \begin{align}
    x &= vt \text{, and because }v=c\text{, we write:} \\
    x &= ct,\label{eq:x=ct} \\
    &= 299792458\text{ ms}^{-1} \times 31557600\text{ s} \\
    &= 9460730472580800\text{ m}, \\
    &= 9460730472580.800\text{ km}.
    \end{align}

    Since light travels this ridiculously large number of kilometres, it makes perfect sense for astrophysicists to use this fact to express the distance of stars and galaxies. This way, the nearest major galaxy, Andromeda, is only about $2.5$ light-years away. This is obviously more practical than $23651826181452\text{ km}$.

    So, about describing distance in terms of units of time, we learnt that

    • in the case of a ‘normal scale’ distance such as between two towns, expressing a spatial distance in units of time makes it easier to compare with the amount of time one wishes to spend on travelling – it becomes like comparing time with time;
    • in the case of larger scale distances such as between two galaxies, expressing a spatial distance in light-units of time makes it easier to handle the impractically large numbers of the original units.

    Let us go back at our diagram in Figure 7 again. The units of both axes are not the same. The $x$-axis is the distance, which is expressed in spatial units, such as metres. The $y$-axis is the time, expressed in temporal units, such as seconds. It is hard to compare the two: the units are not the same. Also, as particle physicists are usually dealing with extremely fast particles, near the speed of light, it is impractical to be using the standard units of time. So, physicists have devised a solution to both problems. Number one: what if we expressed time in units of distance? So, that is the other way around: not distance in units of time, but time in units of distance.

    To do that, we simply use the formula as expressed in Equation \eqref{eq:x=ct}: $x=ct$. In other words, if we multiply time $t$ with the speed of light $c$, we get a distance. A little analysis of units checks out. If we multiply the units of the speed of light with the unit of time, we get a unit of distance:

    \begin{equation}
    \text{m s}^{-1} \times \text{s} = \text{m s}^{-1}\text{s} = \text{m}\frac{\text{s}}{\text{s}} = \text{m}.
    \end{equation}

    Figure 8: ct. (a) The time axis is multiplied by c, so it is easier to compare time with space, i.e. time is expressed in units of distance. (b) We set c = 1 so that the so-called world line of any particle travelling at exactly the speed of light is always at an angle of π/4 with the x-axis. Or 45°, if you are into that sort of thing.

    This does not mean we magically, qualitatively, or even hypothetically transformed the time dimension into a space dimension, even though this would be a perfect device for a cool work of science-fiction, but it does mean that we now express time in units of distance. And so, we label the $y$-axis with $ct$ as is shown in Figure 8.

    Now, to tackle the second problem, where physicists work with particles whose motions approach the speed of light at distance scales smaller than an electron in the vicinity of black holes with forces greater than you would ever encounter, it is impractical to work with the ordinary distance and time units. Furthermore, they prefer to choose the units of $ct$ and $x$ such, that a ‘line’ of a photon, e.g. light, travelling through space and time, is always depicted at an angle of $\pi/4$ or $45^\circ$ with the $x$-axis. To do so, they set the speed of light to 1. So, $c=1$. What you get is a diagram as shown in Figure 8(b). Particle $P$ is a photon, thus travelling at the speed of light. So, its ‘line’ is at the exact angle of $\pi/4$ with both the $x$- and $y$-axis, i.e. the $x$ and $ct$, respectively. All the other particles thus travel at a certain ratio of $c$, i.e. a certain ratio of 1.

    All this should tell you enough to figure out how fast a particle would be going if its ‘line’ would be drawn underneath that of $P$, i.e. at an angle smaller than $\pi/4$. And even though you should also be able to figure out if this is at all possible, we will tell you now that this is not possible.

    By the way, the term ‘line’, which we use to describe the path a particle takes through space and time in our diagrams, is called a ‘world line’ as Hermann Minkowski would have wanted us to. And the diagrams of Figures ref7 and 8 are called Minkowski diagrams. They are also called spacetime diagrams, although there is a subtle difference: Minkowski diagrams are the subset of two-dimensional diagrams within the larger set of spacetime diagrams, which contains the 3D versions, and 4D, even.

    Wick rotation revisited

    Figure 9: from it to ict. (a) The Wick rotation of the real time axis ct to the imaginary time axis ict. We projected the coordinate system of Figure 8 onto ‘the floor’ to have the ct-axis then rotated to the imaginary ict-axis by multiplication by the imaginary unit i. (b) Consequently, the world line of P gets rotated onto the imaginary plane as well.

    We are almost ready to derive the Lorentz transformations. The only thing we have to do, is Wick rotate the (real) time axis into the imaginary time axis, i.e. we rotate the $ct$-axis of Figure 8. So, we do as we did in the previous section Number sets: we multiply by the imaginary unit $i$ from the number set $\mathbb{C}$, thereby rotating the time axis of $\mathbb{R}$ into the complex plane $\mathbb{C}$ to become an imaginary axis of time.

    Figure 9 offers a geometric representation of the whole operation. We projected our original coordinate system of Figure 8 onto ‘the floor’, so to speak. We left out particles $Q$, $R$, and $S$ to keep it legible. Wick-rotating the real time axis $ct$ by multiplying by the imaginary unit $i$ then yields the imaginary time axis $ict$. Automatically, the world line of $P$ rotates along into the complex plane. Note, that the spacetime coordinates of $P$ have changed a few times in this section. They were $(x_P,t_1)$, then they became $(x_P,ct_1)$, and have ended up to become $(x_P,ict_1)$. Just the way we like it.

    Deriving the Lorentz transformations

    Invariant world line in the complex plane

    Figure 10: Wick-rotated M and E. Hermann’s M and Albert’s E frames of reference rotated at an angle θ relative to each other in the complex plane about their origin. We put a little square with sides marked I and II to aid us in our trigonometric calculations.

    To end up with the Lorentz transformations as formulated in Equations \eqref{eq:Lorentz t-prime} and \eqref{eq:Lorentz x-prime} by rotating two frames of reference relative to each other in the complex plane – with an imaginary time axis – we refer to Figure 10.

    In both frames, those of Hermann and Albert, a photon $P$ travels at speed $c$. By the second postulate of Einstein’s Special Relativity[3], we know that, somehow, the value for $c$, which is chosen to be 1 in our case, is the same in both frames of reference, even though one moves relative to the other, meaning the coordinates between the frames are unequal. In Figure 2, this is represented by $v$. In Figure 10, this is represented by an angle $\theta$.

    We see that the coordinates of $P$ in $\mathcal{M}$ are $(\Delta x, ic\Delta t)$. In $\mathcal{E}$, they are $(\Delta x’, ic\Delta t’)$. They are related to each other by some proportion of angle $\theta$. Before we find that relation, we repeat our finding regarding Equation \eqref{eq:interval} in the section Invariances: the quantity $(\Delta x)^2 – (c\Delta t)^2$ is invariant. In our case, it is the interval $OP$ that is invariant, despite the fact that $P$ has different coordinates. In other words, geometrically, both $\mathcal{M}$ and $\mathcal{E}$ agree on the length of the yellow world line as you can see in Figure 10. We should proceed to show this.

    Let us first write down the expressions for the invariant yellow world line $OP$ in both frames of reference:

    \begin{align}
    \text{Hermann, standing in }\mathcal{M}\text{, says: }(OP)^2 &= (\Delta x)^2 + (ic\Delta t)^2, \\
    \text{Albert, standing in }\mathcal{E}\text{, says: }(OP)^2 &= (\Delta x’)^2 + (ic\Delta t’)^2.
    \end{align}

    And, since both expressions calculate the same invariant quantity, obviously, we can write:

    \begin{equation}
    (\Delta x)^2 + (ic\Delta t)^2 = (\Delta x’)^2 + (ic\Delta t’)^2,
    \end{equation}

    which simplifies to

    \begin{equation}
    \Delta x^2 – c^2\Delta t^2 = (\Delta x’)^2 – c^2(\Delta t’)^2.\label{eq:interval in the complex plane}
    \end{equation}

    This is the result we wanted. Whether it is with imaginary time or with real time, the quantity $\Delta x^2 – c^2\Delta t^2$ remains invariant. (Recall that $i^2=(sqrt{-1})^2=-1$.) Even in the complex plane, $\mathcal{M}$ and $\mathcal{E}$ agree on the magnitude of this quantity.

    Coordinates in terms of the other coordinates

    Let us now express the coordinates of $P$ in $\mathcal{E}$, i.e. $(\Delta x’,ic\Delta t’)$, in terms of angle $\theta$ and the coordinates of $P$ in $\mathcal{M}$, i.e. $(\Delta x,ic\Delta t)$. Firstly, we deduce an expression for $\Delta x’$ using Figure 10:

    \begin{align}
    \Delta x’ &= \Delta x\cos\theta + \text{I}, \\
    \text{I} &= ic\Delta t\sin\theta, \\
    \therefore \Delta x’ &= \Delta x\cos\theta + ic\Delta t\sin\theta.\label{eq:delta x prime}
    \end{align}

    Secondly, we deduce an expression for $ic\Delta t’$:

    \begin{align}
    ic\Delta t’ &= ic\Delta t\cos\theta – \text{II}, \\
    \text{II} &= \Delta x\sin\theta, \\
    \therefore ic\Delta t’ &= ic\Delta t\cos\theta – \Delta x\sin\theta.\label{eq:icdelta t prime}
    \end{align}

    Lastly, as we want to find the relation between $\theta$ in the complex plane and $v$ in real spacetime, we forget $P$ for a moment and now write the expression for Albert himself, sitting in $O’$ of his frame $\mathcal{E}$ in terms of the coordinates of Hermann’s frame $\mathcal{M}$. In other words, how does Hermann see Albert move? Since Albert is not moving in his own frame $\mathcal{E}$, as we said earlier, after a certain amount of time $ic\Delta t$, his $\Delta x’=0$. So, by Equation \eqref{eq:delta x prime}, we write

    \begin{equation}
    \Delta x’ = \Delta x\cos\theta + ic\Delta t\sin\theta = 0.
    \end{equation}

    Working this further, we get

    \begin{align}
    ic\Delta t\sin\theta &= -\Delta x\cos\theta, \\
    \frac{sin\theta}{\cos\theta} &= -\frac{\Delta x}{ic\Delta t}, \\
    \tan\theta &= -\frac{1}{ic}\frac{\Delta x}{\Delta t}, \\
    \tan\theta &= -\frac{1}{ic}v, \\
    \tan\theta &= -\frac{v}{ic}.
    \end{align}

    To remove the imaginary unit – being a surd – from of the denominator, we multiply the right hand side with $i/i$, yielding:

    \begin{align}
    \tan\theta &= -\frac{i}{i}\frac{v}{ic}, \\
    \tan\theta &= -i\frac{v}{-c}, \\
    therefore \tan\theta &= \frac{iv}{c}.\label{eq:tan \theta}
    \end{align}

    To recapitulate, we have now obtained Equations \eqref{eq:delta x prime}, \eqref{eq:icdelta t prime}, which express the coordinates of $P$ in $\mathcal{E}$ in terms of angle $\theta$ and the coordinates of $\mathcal{M}$. Lastly, we obtained relation \eqref{eq:tan \theta} between angle $\theta$ and speed $v$ of Albert’s frame $\mathcal{E}$ as seen by Hermann in his frame $\mathcal{M}$. So, to restate, we obtained the following transformations:

    \begin{aligned}\Delta x’ &= \Delta x\cos\theta + ic\Delta t\sin\theta,&\quad\eqref{eq:delta x prime} \\ ic\Delta t’ &= ic\Delta t\cos\theta – \Delta x\sin\theta,&\quad\eqref{eq:icdelta t prime} \\tan\theta &= \frac{iv}{c}.&\quad\eqref{eq:tan \theta}\end{aligned}

    Figure 11: Triangle tan θ. The geometric representation of Equation 9: an imaginary triangle with an imaginary slope tan θ, where Γ is the hypotenuse. Note that sin θ = (iv/c)/Γ and cos θ = 1/Γ.

    The Lorentz transformations

    Note that, algebraically, it is possible to write Equation \eqref{eq:tan \theta} as

    \begin{equation}
    \tan\theta = \frac{iv/c}{1},
    \end{equation}

    which, geometrically, looks like 11. Note that $\sin\theta=(\mathrm{iv/c})/\Gamma$ and $\cos\theta=1/\Gamma$, so all we have to do now, is figure out what $\Gamma$ is. Using, again, the Pythagorean theorem:

    \begin{align}
    \Gamma^2 &= 1^2 + \left(\frac{iv}{c}\right)^2, \\
    &= 1 + \frac{-v^2}{c^2}, \\
    \therefore \Gamma &= \sqrt{1-\frac{v^2}{c^2}}.
    \end{align}

    We can now write:

    \begin{align}
    \sin\theta &= \frac{iv/c}{\sqrt{1-v^2/c^2}}, \\
    \cos\theta &= \frac{1}{\sqrt{1-v^2/c^2}}.
    \end{align}

    This is starting to look good. Moving on to substitute $\sin\theta$ and $\cos\theta$ in Equation \eqref{eq:delta x prime}, yields:

    \begin{align}
    \Delta x’ &= \Delta x \left(\frac{1}{\sqrt{1-v^2/c^2}}\right) + ic\Delta t\left(\frac{iv/c}{\sqrt{1-v^2/c^2}}\right), \\
    &= \frac{\Delta x}{\sqrt{1-v^2/c^2}} + \frac{-v\Delta t}{\sqrt{1-v^2/c^2}}, \\
    \therefore \Delta x’ &= \frac{\Delta x-v\Delta t}{\sqrt{1-v^2/c^2}}.
    \end{align}

    Since in our configuration the differences are calculated from the origin, we can leave out the $\Delta$-sign, using just the coordinates, and so we obtain

    \begin{equation}
    x’ = \frac{x-vt}{\sqrt{1-v^2/c^2}},
    \end{equation}

    which is indeed Equation \eqref{eq:Lorentz x-prime}.

    Substituting $\sin\theta$ and $\cos\theta$ in Equation \eqref{eq:icdelta t prime}, yields:
    \begin{align}
    ic\Delta t’ &= ic\Delta t\left(\frac{1}{\sqrt{1-v^2/c^2}}\right) – \Delta x\left(\frac{iv/c}{\sqrt{1-v^2/c^2}}\right), \\
    ic\Delta t’ &= \frac{ic\Delta t}{\sqrt{1-v^2/c^2}} – \frac{iv\Delta x/c}{\sqrt{1-v^2/c^2}}, \\
    \Delta t’ &= \frac{\Delta t}{\sqrt{1-v^2/c^2}} – \frac{v\Delta x/c^2}{\sqrt{1-v^2/c^2}}, \\
    \therefore \Delta t’ &= \frac{\Delta t-v\Delta x/c^2}{\sqrt{1-v^2/c^2}}.\end{align}

    And so, leaving out the $\Delta$-sign, using just the coordinates, we obtain
    \begin{equation}
    t’ = \frac{t-vx/c^2}{\sqrt{1-v^2/c^2}},
    \end{equation}

    which is, indeed, Equation \eqref{eq:Lorentz t-prime}.

    It is important to note that, while not unusual to leave out the $\Delta$-sign, formally, it is incorrect: in Special Relativity there is no preferred (fixed) origin, hence, it is always about differences.

    Lastly, we reiterate that the term $1/\sqrt{1-v^2/c^2}$ is often written as $\gamma$ and is called the Lorentz factor. Also, in some texts, the term $v/c$ is replaced by symbol $\beta$, yielding the following equivalent expressions of the Lorentz transformations:

    \begin{align}
    ct’ &= \gamma(ct-\beta x), \\
    x’ &= \gamma(x-\beta ct), \\
    y’ &= y, \\
    z’ &= z.
    \end{align}

    Thanks to the imagination of many mathematicians and physicists before us, our ability to investigate, analyse, and calculate has become as supple and malleable as is, indeed, the fabric of the cosmos.

    Featured image: arielrobin

    [1] Hawking, S. (2001) The universe in a nutshell. New York: Bantam Books.
    [2] Walter, S. (2014) Poincaré on clocks in motion. Amsterdam, Ne.
    [3] Einstein, A. (1905) “Zur Elektrodynamik Bewegter Körper,” Annalen der Physik, 322(10), pp. 891–921. doi: 10.1002/andp.19053221004.
    [4] Bailey, D. H. and Borwein, J. M. (2016) Pi : the next generation : a sourcebook on the recent history of pi and its computation. Switzerland: Springer. doi: 10.1007/978-3-319-32377-0.
    [5] Cooke, R. (2005) The history of mathematics : a brief course. 2nd edn. New York, N.Y.: Wiley.
    [6] Caparrini S. (2006) On the Common Origin of Some of the Works on the Geometrical Interpretation of Complex Numbers. In: Williams K. (eds) Two Cultures. Birkhäuser Basel, pp. 139-151.
    [7] Wick, G.C. (1954) Properties of Bethe-Salpeter Wave Functions. Physical Review, 96(4), pp. 1124-1134.