Author: KJ Runia

  • Gravity Assist: the Planetary Slingshot

    Gravity Assist: the Planetary Slingshot


    Last Christmas, my father and one of my brothers wondered how a spacecraft’s planetary slingshot works. It’s a well-known manoeuvre to get a big swing forward by passing by a planet. They realized that if the craft falls toward the planet due to its gravitational pull, it will gain momentum. However, as soon as it has flung around the planet, wouldn’t that same gravitation slow down the craft just as much? Great question, of course, so, let’s dive into this gravity assist short and simple.


    Gravity assist

    First things first, you are absolutely right in thinking that the gravitational pull of the planet is a symmetrical situation. The same ‘force’ pulling the craft toward the planet making it accelerate, will also slow it down once it has reached its maximum point around the planet when it tries to escape the planet’s influence again, continuing its journey in space.

    So, the whole planet’s gravity system is entirely symmetrical. What isn’t symmetrical thoughโ€ฆ is that the planet has an orbital momentum(beginfootnote)Momentum is mass times velocity in a specific direction(endfootnote) around the Sun! The planets in our solar system all orbit around the Sun. Though at different rates, their position relative to the Sun changes every single second. This means they have a great deal of momentum and we, the rather clever clumps of human cells that we sometimes are, can exploit this. So, by using the gravitational pull of the planet to get closer to it, the craft gains an additional tug by the simple fact that the planet itself is also moving through space. And so, while the momentum the spacecraft gained by the planet’s gravitational pull was lost again while escaping the planet’s gravitation, it did gain a bit of the planet’s orbital momentum!

    It’s like catching a train while it’s hurling past the platform. Suppose you’re Batman and you possess a grappling hook. As soon as the train is coming through the station, you start running along with it on the platform. You pull out your grappling hook gun and grapple onto the train. Then you press the button so that the retractable cable attached to the hook pulls you toward the train (still speeding along its trajectory). At the right moment, you release the hook, spread out your batwings, and away you go, having parasitized off of the train’s momentum for a little bit.

    And you’re right, this also means that the train lost a bit of its momentum due to you ‘pushing off of it’. However, due to the huge difference between your mass and the train’s mass, in combination with either velocities, the train’s momentum loss is negligible while your momentum increase is significant.

    The same goes for the planet. A spacecraft doing a gravity slingshot, or, enjoying a gravity assist, as it’s also called, will decrease the orbital velocity of the planet, meaning that the planet’s orbit will get closer to the Sun. Of course, given the ratio between the spacecraft’s mass and the planet’s mass, this amount is negligible.

    Voyager

    Two of the most famous instances of gravity assists are the voyages of Voyager 1 and 2. In the animation below, you can see the trajectory of Voyager 2, where it gained assists from several of our solar system’s planets. Earth is the fast-orbiting blue dot. Jupiter is green, Saturn is cyan, Uranus is yellow, and orange is Neptune. And while the Voyager spacecrafts flew by the planets, they sent some of the best postcards back to Earth.

    Made by Phoenix7777, published under CC BY-SA 4.0

    Deceleration

    Of course, what goes around, may come around too. If you’d like the spacecraft to decelerate, you make it fly opposite the planet’s orbital motion. This way it slows down while donating a (negligible) bit to a planet’s momentum.

    So, what Christmas message can we take from this? That’s right: even the tiniest entity in the Universe is capable of changing an entire planet’s momentum. No matter how small, in one way or another, its effects are significant.

  • Why is glass transparent?

    Why is glass transparent?


    Imagine transparent materials didn’t exist. What would cars look like? How would you be able to look at the cold Winter Moon from your bedroom window without getting cold yourself? Would air be opaque too? What about the lenses in our eyes? And if all materials consist of molecules, how is it that the molecules of glass are transparent while others aren’t? This last question was asked by the oldest(beginfootnote)If I’m not mistaken, thirteen at the time I’m writing this.(endfootnote) son of one of my best friends last week.

    Firstly, I’ll try to give an answer as short as possible. If you’d like to know more, you can read on. Be warned, however: the article is quite possibly a bit long. It’s just that the answer to this seemingly easy question asks for quite some background knowledge. On the other hand, with a brilliant question like this oneโ€”let’s just say, you’re asking for it.

    Short answer

    Sometimes the composition of a molecule is such that its electrons will hardly respond to passing photons. For all intents and purposes, they will leave them be. At most, they will change their course a little. To our eyes, the material is then transparent.

    If electrons do react to incoming photons, they might reflect them or absorb all of their energy, making the photons disappear. Sometimes they just vibrate a little bit and make the entire atoms vibrate a little bit too (phonon), but nothing much else happens. The material just increases temperature for a minuscule amount. Sometimes the electrons vibrate so much that they’ll radiate that energy very soon after, causing new photons to be created, which then move on to the rest of the universe. To us, that material is then opaque; the photons radiated from the material end up in our eyes.

    So, this was the short version. If you’d like to know more, do read on!

    Molecules, atoms, elementary particles

    Perhaps you know this but just to be sure: all solids, liquids, and gasses consist of molecules. Here you see a photo of a bunch of so-called pentacene molecules made by Canadian scientists[1]. These molecules aren’t present in glass, however, it does give an impression of what molecules can look like.

    caterpillar-like pentacene molecules

    Every caterpillar-like thingy is a molecule. When they stick close enough together, they form a solid. If they are capable of sliding past each other, it’s a liquid. And if they’re capable of jiggling a lot more away from each other, it’s a gas.

    It’s possible to look more closely. Those molecules consist of atoms. Do have a look at this cool photo of one such molecule which Swiss physicists and one physicist at Utrecht University were able to snap in 2009[2].

    a pentacene molecule, consisting of five benzene rings

    You might be able to distinguish five hexagonal shapes with protrusions. At every corner and every protrusion, an atom is present. You don’t see the individual atoms โ€“ they’re too small for that. However, you can see the structure formed by the chain of atoms, thereby shaping the molecule into existence.

    Back to glass. The window in your bedroom is composed of different types of molecules. There are a lot of silicon dioxide molecules, sodium carbonate molecules, calcium oxide molecules, magnesium oxide molecules, and some aluminium oxide molecules. Here you see a drawing of one such silicon dioxide molecule.

    a molecule of silicon dioxide

    Do they look like that for real? No, absolutely not. It’s just a conceptual model. In science, a model is meant to be a tool and never an exact copy of reality. And yet, we use models as they are quite helpful for imagining what we’re working with and for doing calculations on them. You have to keep in mind, though, it’s not what it really looks like.

    From (the model of) the silicon dioxide molecule you can see that it is comprised of three atoms: one silicon atom (grey) and two oxygen atoms(beginfootnote)โ€˜Oxygenโ€™ and โ€˜oxideโ€™ stem from the ancient-Greek words for โ€˜sharpโ€™ (แฝ€ฮพฯฯ‚, oxรบs), and โ€˜birthโ€™ (ฮณฮญฮฝฮฟฯ‚, gรฉnos), and the Latin word for โ€˜acidโ€™, which is acidus. Lastly, the word โ€˜diโ€™ stems from the ancient-Greek ฮดฮฏฯ‚ (dรญs), meaning โ€˜twiceโ€™. As the molecule consists of two oxygen atoms, the official chemical name of the molecule is thus silicon dioxide.(endfootnote) (red). You may also wonder what these two little bars on each side are supposed to be. They symbolise the electrons which are shared by all the atoms amongst each other. Atoms can stick together when they share electrons with each other. In other words, it refers to how strong the atomic bond is. More bars equals a stronger bond.

    Every atom consists of yet smaller parts. Apart from one(beginfootnote)the hydrogen atom(endfootnote), atoms are made up of three types of particles: electrons, protons, and neutrons. At the core of the atom are all the protons and neutrons. The electrons kind of swirl around them in a cloud-like type of existence. As electrons don’t themselves consist of smaller things, they are said to be elementary(beginfootnote)โ€˜Elementaryโ€™ stems from the Latin word elementum, carrying a meaning like โ€˜first principleโ€™. There is nothing that goes further down than what is elementary. Elements always form the basis for other things.(endfootnote) particles. Here you see a model of an atom.

    a model of an atom

    The core (or โ€˜nucleusโ€™) with all the protons and neutrons is so small that it’s usually drawn as a point or a little ball. However, if you were to zoom in, you’d see a lump of protons and neutrons. Around it there is the cloud-like electron or multiple electrons. (If you’d like to know exactly why a cloud is the model for one or more electrons, you can read This is not an atom.)

    It’s quite possible that an electron is further removed from the nucleus than shown here. If an electron receives energy, it’ll jump further away from the nucleus. After a very short period, the electron might jump back to its old position. If it does so, its energy leaves the atom again in the form of light. One of the ways in which the electron receives energy is light.

    Light

    What is light? Is it a wave, does it consist of particles? For a long time, physicists had no idea what light was exactly. Since the seventeenth century, great debates went on between physicists supporting Sir Isaac Newton and physicists supporting Christiaan Huygens. I fear a little that some science teachers at high schools still think it’s a big mystery. One of my science teachers in high school told us he was still on the fence whether it’s particles or waves. Unfortunately for him, since slightly less than a hundred years ago, we know.

    Isaac Newton (left): light = particles (tiny balls). Christiaan Huygens (right): light = waves. The correct quantum mechanical answer is: light = a disturbance in the Universe-pervading electromagnetic field. Depending on what is practical, one uses the mathematics of classical waves or the mathematics of photons (wave packets, not balls!) to work with. Physics students learn to use both approaches.

    What I’m about to tell you is not something you’ll likely learn in high school. I’m not sure why but it may have to do with textbook authors finding the mathematics too complicated. So, what you’re about to read is more or less what you’ll learn at university as a physics student, only without the mathematics.

    The problem with the question โ€˜wave or particleโ€™ is that it suggests there’s only one choice. This is incorrect. The question should be: What is light? The answer is: a field(beginfootnote)In mathematics and physics, we call this a gauge field. It’s quite abstract mathematics. However, no matter how abstract, it has proven to be highly applicable in practice. Mobile phones would not have existed without these abstract mathematics.(endfootnote), one of the many Universe-pervading fields present, in this case the electromagnetic field. And to be even more precise: light is a disturbance of this electromagnetic field. One can describe this disturbance as either a wave or a particle, depending on what is more practical for the matter at hand.

    Besides, in modern physics the meaning of the word โ€˜particleโ€™ differs from what you’d normally expect. In physics a particle is actually a packet, a wave packet. It’s not a pellet, it’s not a tiny ball or even a point. It’s a tiny packet of information which we mathematically describe as a tiny wave (a disturbance).

    Richard Feynman was an important, Nobel Prize-winning physicist who made enormous contributions to quantum electrodynamics.

    According to one of the best theories we have of our Universe to date, so-called quantum electrodynamics(beginfootnote)โ€˜Quantumโ€™ is Latin for โ€˜how muchโ€™. Physicists have been using the word as a synonym for โ€˜particleโ€™. Plural is quanta. โ€˜Dynamicsโ€™ stems from the ancient-Greek ฮดฯ…ฮฝฮฑฮผฮนฮบฯŒฯ‚, dunamikรณs, โ€˜powerfulโ€™ en refers to the theory describing forces and change of forces.(endfootnote) (QED), the Universe is pervaded by a mostly invisible โ€“ yet sometimes visible! โ€“ electromagnetic field. In most cases, that field does nothing at all. You can’t smell it, you can’t touch it, you can’t see it.

    However, when the electromagnetic field is being disturbed at a specific place in the Universe โ€“ e.g. on the inside of the LED lamp in your lavatory โ€“ then that disturbance will propagate in all directions, from that specific spot in the Universe towards the very rest of the Universe โ€“ i.e. the space of your lavatory. This disturbance you can see! I’m not sure how your pets might call this disturbance, however, humans call it light.

    Albert Einstein
    Albert Einstein

    If your eyes were able to zoom in immensely, you would see that light actually consists of billions and billions and billions of tiny disturbances. Light is a bundle of tiny disturbances in the omnipresent electromagnetic field. Those tiny disturbances used to be called โ€˜light quantaโ€™ by Albert Einstein and others. However, since 1928, we call them photons(beginfootnote)This stems from the ancient-Greek ฯ†แฟถฯ‚, phรดs, which ironically means โ€˜lightโ€™.(endfootnote). In popular books and magazines and even by physicists they are called โ€˜particlesโ€™. Again, they’re not pellets or tiny balls or anything. The word โ€˜particleโ€™ refers to them being very tiny but it doesn’t say anything about what they look like. As model, tiny pellets or points are sometimes used, however, it’s not what they are. Photons are, just like electrons, elementary, however.

    In high school and at university, to do calculations on light, the wave model of light is used rather often. The great mathematician and physicist James Clerk Maxwell was one of the founders of the mathematical framework of the wave model of light. He and others before him are responsible for us still talking about โ€˜light wavesโ€™ instead of photons. The classical electromagnetic theory of Maxwell works so well that it’s compulsory for physics students to study this wonderful theory. So, it’s not at all wrong to speak of light waves.

    James Clerk Maxwell
    James Clerk Maxwell

    In the twentieth century, however, physicists found that quantum electrodynamics was able to predict and describe more phenomena than Maxwell’s classical electromagnetic theory, so the first kind of replaced the latter. Put differently, Maxwell’s theory is still highly useful in industrial applications, however, with QED, you can do what Maxwell’s theory can do plus a lot more.

    This is the way it usually goes in physics. The law of universal gravitation of Sir Isaac Newton works brilliantly. You can even apply it to Mars landings. However, the theory of gravity by Albert Einstein, so-called general relativity, can do what Newton’s theory does and a lot more, more precisely. So, general relativity has kind of replaced Newton’s law of universal gravitation. And yet, the latter is compulsory in high school and at university. It’s not wrong. It’s very useful, even! However, it does have its limitations. That’s why we first learn about Newton’s gravity and only later do physics students have to learn about Einstein’s gravity. Without Einstein’s general relativity, Google Maps and GPS-systems inside cars would not have worked properly.

    Hence, physics students learn everything about Maxwell’s wave theory and only later do they learn about quantum electrodynamics. And without quantum electrodynamics you would not have had computer processors, there would have been no internet, no mobile phones, no touchscreens.

    Light and energy

    The great physicist Max Planck came up with the idea that every photon has a specific energy level. He also showed that with every energy level comes a particular light colour. Bright blue light carries more energy than deep-dark red light. Sometimes light (photons) has (have) so much energy that it has (they have) become invisible to our human eyes. High-energy ultraviolet(beginfootnote)โ€˜Ultraโ€™ is Latin for โ€˜beyondโ€™. So, ultraviolet means beyond violet.(endfootnote) light (UV light) is invisible to us. However, if your eyes were much more sensitive than they are now, you would see a very bright โ€˜more violet than violet-colouredโ€™ light. Conversely, light can have very little energy. So little even, we won’t be able to see it anymore. Hence, infrared(beginfootnote)โ€˜Infraโ€™ is Latin for โ€˜belowโ€™. So, infrared is โ€˜belowโ€™ or โ€˜less thanโ€™ red.(endfootnote) light is invisible to us. However, if our eyes were slightly more sensitive, we would see โ€˜less red than red-colouredโ€™ light.

    This cheery looking fellow was a physicist and a genius. His name was Max Planck. A photograph from 1933.
    This cheery looking fellow was a physicist and a genius. His name was Max Planck. A photograph from 1933.

    WiFi and 4/5G are light too. The photons have very little energy compared to the photons in your lavatory. If our eyes had been thousands of times more sensitive than they are now, you would have seen that the antennas of the WiFi router and the mobile phones are basically lamps radiating โ€˜less than less than less than (thousands of times โ€˜less thanโ€™) red-colouredโ€™ light.

    The electromagnetic field pervading our Universe can thus be disturbed at various energy levels. Depending on that, light looks differently. It has varying colours or is invisible โ€“ which it is most of the time. Our eyes aren’t the best instruments to look around with. Of all possible energy levels the electromagnetic field can be at, we can only discern just a few. That energy portion is what we call visible light.

    Another word for disturbances of the electromagnetic field is electromagnetic radiation. Depending on the energy level of the radiation, we have different terms for it, such as โ€˜radioactive radiationโ€™ or โ€˜gamma radiationโ€™. However, all these things โ€“ the lavatory light, the WiFi, 4/5G for the mobile phone, the head lights of the car, the radio waves from the neighbour, the Bluetooth speaker in the kitchen, the x-ray images at the dentist, the microwave โ€“ are all light, are all electromagnetic radiation, are all disturbances of the one and the same electromagnetic field. The only difference is the energy level of that disturbance.

    The correct order going from very little to deadly amounts of energy is the following: radio, WiFi, microwave(beginfootnote)If you want to know whether microwave radiation is deadly or not, do give my article Is microwave oven radiation unhealthy? a read.(endfootnote), 4/5G >> infrared light (TV remote) >> visible light (lavatory light, club lights) >> UV light (take care, apply sunscreen) >> x-rays (only operated by professional medical workers) >> gamma radiation (deadly, except for Bruce Banner) >> cosmic radiation (deadly, except for Captain Marvel).

    As most electrons inside of walls of houses won’t respond much to electromagnetic disturbances (photons) at the energy level of WiFi (very little energy), to WiFi photons, the walls are almost transparent. This is why you can receive WiFi straight through the walls. If your eyes were sensitive enough, you would be able to see the light coming from the router, straight through the walls. Those same electrons, however, do react to photons at the much higher energy level corresponding to visible light. This is why those photons do not fly through the wall. And this is why we find walls to be quite the opaque type objects. Nevertheless, the electrons do not respond again to photons at the even higher โ€“ much higher โ€“ energy levels of x-rays. This is exactly why walls are perfectly transparent to Superman.

    Depending on the composition of the molecules and atoms do electrons more or less react to the presence of photons at varying energy levels. If electrons of the material do not respond to photons at the energy level corresponding to visible light, then the material is transparent to us.

    Below you see a diagram of the full spectrum(beginfootnote)โ€˜Spectrumโ€™ is Latin for โ€˜appearanceโ€™. So, if you speak of the spectrum of something, such as electromagnetism, then you’re referring to all of its appearances.(endfootnote) of electromagnetic radiation (click to enlarge). As you can see, only a small portion is visible to us.

    A diagram of electromagnetic radiation. Far right, we see the dangerous types of radiation: cosmic rays, x-rays, gamma rays, UV-light. In the middle, we see visible light. Far left, we see the lowest energy photons: WiFi, mobile phones, microwave ovens.
    A diagram (not to scale) of electromagnetic radiation, or photons, if you will. The mentioned values are the frequencies of the photons, expressed in gigahertz (GHz). The higher the frequency, the higher the energy of the photon.

    GHz refers to the frequency of the photon and is a measure for the photon’s energy level. The higher the frequency, the higher the energy level. Ultraviolet radiation is where it’s starting to become dangerous to us. This is where our cells become damaged (โ€˜DNA damageโ€™). As long as you’re not exposed to the Sun for too long and x-ray photography is done in very short amounts of time, it’s going to be fine. But be careful. Again, gamma radiation and cosmic radiation are deadly. I understand, it doesn’t feel comfortable at all, but please, please do listen to your parents when you’re going out for a space walk. Put on that spacesuit.

    Impressionable electrons

    In the previous century, physicists such as Albert Einstein discovered that electrons can be influenced by incoming photons(beginfootnote)This is what he received the Nobel Prize for. You can read more about that in my article The formula that got Albert Einstein the Nobel Prize and should stop us getting sunburn all the time.(endfootnote). It very much depends on the way the electrons are captured inside the molecules โ€“ which depends on the type of atoms โ€“ at which energy level photons they will start reacting.

    In the case of glass, the electrons do feel electromagnetic disturbances slightly. This is why they do start to jiggle differently just a notch. That jiggling causes changes in the part of the electromagnetic field that is inside of the glass. And these changes will influence the photons (disturbances in that same electromagnetic field) in such a way that they’ll change course slightly.

    There's a bear swimming in a pool in a zoo. The pool is visible from the side through a large window. Due to refraction, the head of the bear above the surface seems to be located at a different place than the rest of its submerged body. The bear seems beheaded and yet, it lives.

    This is why the image behind glass can seem to be slightly warped. The same happens when light goes from air to water (as water, too, contains electrons which react to incoming photons). The electrons don’t do too much so that photons can just pass through, however, they do enough so that the photons do change course slightly. Or a lot as you can see by the water in the photo above. In my article Why, exactly, do glass and liquids refract light? we take a deep dive into this phenomenon.

    Why is glass transparent?

    And so, glass is transparent as the electrons in glass molecules aren’t capable of reacting very much to incoming electromagnetic disturbances (photons). Just a little. So, they do bend the original trajectory of the photons slightly.

    There are materials containing electrons responding to all energy levels except those corresponding to blue light, for example. This means that blue light can just pass through while the rest is being absorbed. To us, this material seems to be a blue filter.

    It’s also possible to produce materials carrying electrons which react to all photons in the visible part of the spectrum. They do this so strongly that photons will be reflected completely. We call that a mirror.

    closeup photo of primate looking in a mirror

    Note that we’re talking mostly about photons we can see. To us most glass is transparent. However, we can also produce glass which seems transparent as it lets visible light pass through, while they are much less transparent to birds at the same time.

    We are incapable of seeing UV light. Birds can, however. So, if glass is produced in such a way that they will let visible light pass through but not UV light, they seem less transparent to birds, preventing them to bump into it.

    So, the answer to the question, โ€˜And if all materials consist of molecules, how is it that the molecules of glass are transparent while others aren’t?โ€™, should rather be: โ€˜Transparent to whom? To birds? Or to humans?โ€™

    References

    [1] Dinca, L. E. et al. (2015) โ€œPentacene on Ni(111): Room-Temperature Molecular Packing and Temperature-Activated Conversion to Graphene,โ€ Nanoscale, 7(7), pp. 3263โ€“3269. doi: 10.1039/C4NR07057G.

    [2] Gross, L. et al. (2009) โ€œThe Chemical Structure of a Molecule Resolved by Atomic Force Microscopy,โ€ Science, 325(5944), pp. 1110โ€“1114. doi: 10.1126/science.1176210.

  • The Copenhagen interpretation

    The Copenhagen interpretation


    So, you’re reading this post on a quantum mechanical thing, be it a portable device or a less mobile desktop computer. This is all possible because quantum mechanics is the most successful theory humans have been able to conjure up. Nevertheless, there are still a few things to figure out. One of those things is called the measurement problem and sits at the level of a Nobel Prize. Since 1925, one way of solving this was mainly proposed by Niels Bohr and Werner Heisenberg. They coined what is now called the Copenhagen interpretation.

    Wave function

    Let’s do a quick recap of the mechanics as described in The double-slit experiment. Suppose, we release a bunch of free electrons, meaning that they are not disturbed nor confined by any interaction with any other thing. They are launched from a cannon towards a screen with two slits.

    One important element to describe the behaviour of the electron is what is called the wave function (the other element is the Schrรถdinger equation). It’s a mathematical description of all the possible states the electron can be in as soon as you measure it. In the case of our double-slit experiment, we look at two possible states pertaining to its position: it can be in the position-state of being at slit 1 or it can be in the position-state of being at slit 2.

    Figure 1. The screen with slits 1 and 2.
    Figure 1. The screen with slits $\lvert 1 \rangle$ and $\lvert 2 \rangle.$

    Let’s use the symbol $\lvert \Psi \rangle$ to denote the wave function of the electron. Indeed, this is the actual notation for the wave function in quantum mechanics. We’ll use $\lvert 1 \rangle$ and $\lvert 2 \rangle$ to denote the position-states slit 1 and slit 2.

    So, when we don’t put particle-measuring detectors at both slits, we will not know through which slit the electron will have gone. Quantum mechanics dictates that the wave function of our free electron with respect to its position in space is the addition of the two possible slits it can go through. It is true that it could also bump into the screen, missing the slits. So, let’s label all those positions $\lvert b_n \rangle,$ where $b$ stands for bumping-into-screen and $n$ is a number denoting every tiny position on the screen that is not a slit. The wave function can then be written(beginfootnote)I’m deliberately leaving out any complex coefficients so as to not complicate things for the purpose of this post.(endfootnote) as follows:

    $\lvert \Psi \rangle = \lvert 1 \rangle + \lvert 2 \rangle + \lvert b_1 \rangle + \lvert b_2 \rangle + \dots + \lvert b_n \rangle$

    Put differently, we say that with respect to its position, the electron is in a superposition of slit 1 and slit 2 and a whole bunch of other positions on the first screen that lead to nowhere but an inglorious end on the first screen.

    The unmeasured, free particle is wave-like โ€“ going through the two slits at once, spread-out like a wave โ€“ as is visible on the second screen: after a while a wave-interference pattern will have emerged.

    Figure 2. The typical interference pattern for wave-like phenomena becomes visible on the second screen when free electrons go through the two slits.
    Figure 2. The typical interference pattern for wave-like phenomena becomes visible on the second screen when free electrons go through the two slits.

    Measurement

    However, as soon as you attach particle detectors on both slits, the detectors show that the particle only ever goes through one slit. Also, the interference pattern disappears instantly and makes place for your typical particle-like pattern.

    The detectors determine or measure the position of the electron to be either at slit 1 or slit 2. Somehow, they never measure the electron to be at slit 1 and 2 at the same time. In other words, its wave-like existence has been replaced by a particle-like existence!

    Figure 3. The typical particle-like pattern appears on the second screen as soon as you put detectors at the slits to measure the position of the particle.
    Figure 3. The typical particle-like pattern appears on the second screen as soon as you put detectors at the slits to measure the position of the particle.

    The Copenhagen interpretation

    Undergraduate students are usually taught the following interpretation of these strange events.

    Upon measurement, the particle’s wave function collapses.

    That’s it. This is what’s at the heart of the Copenhagen interpretation. This explanation of what’s happening at the act of measurement is named after the city where Niels Bohr worked.

    In other words, all the terms of the wave function โ€˜collapseโ€™ into just one term. Suppose, we measured the electron to be going through slit 2, then by some unknown mathematical operation our wave function

    $\lvert \Psi \rangle =\lvert 1 \rangle + \lvert 2 \rangle + \lvert b_1 \rangle + \lvert b_2 \rangle + \dots + \lvert b_n \rangle$

    gets all the terms cancelled like

    leaving the measurement outcome like

    $\lvert \Psi \rangle = \lvert 2 \rangle.$

    Note that neither the measurement nor the mathematics influence or prescribe at all which terms eventually get cancelled. In this example, it just happens to be that $\lvert 2 \rangle$ was left over. What is eventually being crossed out is fundamentally unpredictable. Moreover, the very process of crossing out is unknown.

    The Born rule

    Max Born wrote in a Nobel Prize-winning footnote of his 1926 paper that the probability of a solution to the Schrรถdinger equation of a quantum-mechanical system (such as an electron) is proportional to the wave function squared. A solution to the Schrรถdinger equation represents a possible specific quantum state, such as being at slit 2. In our equations above, each term represents such a specific quantum state.

    So, the only thing we have is the Born rule(beginfootnote)It would have been too cheesy and/or dorky perhaps, but I would have loved Robert Ludlum’s trilogy to at least contain one with the title The Bourne Rule: One Man Against the Odds.(endfootnote), stating that we can only predict the probability of each of the possible terms to not be crossed out after measurement.

    Figure 4. The footnote that got Max Born the Nobel Prize (reference 1).
    Figure 4. The footnote that got Max Born the Nobel Prize[1].

    Acceptance of the Copenhagen interpretation

    Bohr stated that this collapsing process cannot be described by quantum mechanics. This interpretation seems rather unsatisfactory to many. As the Nobel Prize laureate Steven Weinberg notes, โ€˜This answer is now widely felt to be unacceptableโ€™[2]. A growing number of physicists realise that one of the problems is that in this case nobody knows what is supposed to be described by quantum mechanics and what not. If wave function collapse oughtn’t fall under the purview of the mathematical formalism, for example, then what criteria should we uphold to determine if something else can be studied quantum-mechanically and what cannot be known, ever?

    The major proponents Heisenberg and Bohr stressed that the wave function is purely a mathematical affair. While their views did not always align perfectly, they obviously agreed on wave function collapse[3]. Bohr noted that we had to give up any physical representation of the whole matter. Those who uphold the traditional Copenhagen interpretation are generally instrumentalists.

    However, nobody knows if the Copenhagen interpretation is correct. There are other contenders aiming to solve the measurement problem.

    We will discuss other instrumentalist and realist interpretations of quantum mechanics in another bit of maths and physics.

    Featured image: Werner Heisenberg (left) and Niels Bohr (right) by Fermilab, U.S. Department of Energy. Public domain.

    References

    [1] Born, M. (1926) โ€œZur Quantenmechanik Der StoรŸvorgaฬˆnge,โ€ Zeitschrift fuฬˆr Physik, 37(12), pp. 863โ€“867. doi: 10.1007/BF01397477.

    [2] Weinberg, Steven (2018) โ€œ14 the Trouble with Quantum Mechanics,โ€ in Third Thoughts. Cambridge, Massachusetts; London, England : Harvard University Press, 2018, pp. 124โ€“124.

    [3] Kiefer, C. (2003) โ€œOn the Interpretation of Quantum Theory โ€” from Copenhagen to the Present Day,โ€ in Castell, Lutz and Ischebeck, Otfried (eds.) Time, Quantum, and Information. Berlin; New York : Springer, 2003, pp. 291โ€“299. doi: 10.1007/978-3-662-10557-3_19.

  • Spaces and dimensions

    Spaces and dimensions


    As is usually the case with scientific buzz words in everyday parlance, in books, in the cinema, on TV, and on the internet โ€“ like energy โ€“ the meaning of the word dimension rarely aligns with what mathematicians and physicists understand it to be. This day and age it’s rather uncommon to not have been exposed to phrases such as โ€˜higherโ€™ or โ€˜other dimensionsโ€™. It’s likely you’ve used them yourself once or twice in your life. In this episode, we’ll explore what mathematicians and physicists mean when they talk about dimensions, and, more interestingly, the spaces they yield.

    Dimensions are not Universes

    In science-fiction or even everyday lingo, the word โ€˜dimensionโ€™ is often synonymous with entire worlds, or realms or (pocket) Universes. For instance, aliens may have come from another dimension. Or souls or โ€˜essencesโ€™ dwelling on a โ€˜higher plane of existenceโ€™ in another โ€˜dimension of realityโ€™ are spoken about.

    On a regular basis, portals to other dimensions are opened from which exotic forms of matter and energy are extracted to benefit either the hero or the bad guy of the story.

    It’s also a favourite way to travel great distances within our reality. Just hop through a dimensional portal and out you come, back into our reality, only thousand kilometres away from where you started. Occasionally, you may also travel in time by flying through other dimensions.

    And, of course, other dimensions can be summoned into our own reality or, if the story goes that they have always been present inside our reality, they can be made visible by powerful minds. This is, again, alluding to dimensions being whole separate realms within our realm.

    Figure 1. Doctor Stephen Strange (Benedict Cumberbatch) is about to step into the Mirror Dimension as summoned within (or next to) our reality by his mentor, the Ancient One (Tilda Swinton), in the 2016 film Doctor Strange of the wildly popular Marvel Cinematic Universe (MCU).
    Figure 1. Doctor Stephen Strange (Benedict Cumberbatch) is about to step into the Mirror Dimension as summoned within (or next to) our reality by his mentor, the Ancient One (Tilda Swinton), in the 2016 film Doctor Strange of the wildly popular Marvel Cinematic Universe (MCU). License note. (Click to enlarge.)

    This whole section was just to let you know that what is meant by dimensions in most science-fiction stories is not what is meant in mathematics and physics. They are not realms, realities, worlds or pocket Universes. If we were to refer to realms, realities, worlds, and Universes, we would just say realms, realities, worlds, and Universes, but not dimensions.

    Ordinary spaces and dimensions

    So, what do mathematicians and physicists mean when they talk about dimensions?

    In many cases, they pertain to the actual directions you and I are able to travel in ordinary space. I prefer to think of birds and fish as gorgeous examples of being able to travel in all directions of space all by their own.

    They can fly from your left to your right and vice versa (first direction). They can fly head-on towards you and whizz by over your head and fly further behind you and vice versa (second direction). And, obviously, they can fly up from underneath you and they can keep on flying to way above your face. And vice versa (third direction).

    In many cases, all three directions are oriented perpendicularly with respect to each other. To use another word, they are orthogonal. All motion can be described as some combination of moving in these three orthogonal directions, i.e. orthogonal dimensions.

    In high school we have gotten all too familiar with these three dimensions. We were tortured with finding distances between vertices of a cube along the edges, the sides, and straight through the block. Of course, this is what modern gadgets and cinematography refer to when they use the term 3D, three-dimensional. In some way or form, all three orthogonal dimensions are either taken advantage of or simulated in a virtual way.

    Mathematically, the capability of travelling (or โ€˜transportingโ€™) along these three orthogonal directions automatically give rise to a space, a topology, of some shape or form. Ordinary space is the space you and I are born in and have grown very much accustomed to.

    So, while dimensions may give rise to spaces, they are definitely not the same. Besides, while one dimension by itself technically yields a topology, a space, it’s still a one-dimensional space, meaning, no three-dimensional bodies are able to traverse this without being torn apart.

    Figure 2. In ordinary space, we have three dimensions in the x-direction, the y-direction, and the z-direction. In high school, we were to calculate the distance between points O and F, for instance.
    Figure 2. In ordinary space, we have three dimensions in the $x$-direction, the $y$-direction, and the $z$-direction. In high school, we were to calculate the distance between points $O$ and $F,$ for instance.

    Euclid and Descartes

    A very informal definition of dimensions is the number of coordinates needed to locate an object (in a space of some kind). So, on a flat surface (a plane), such as a ceiling, you need two coordinates to locate a fly. A fly can be 2 metres away from the left wall (the first direction) and 3 metres away from the back wall (the second direction, perpendicular to the first direction). Its coordinates are therefore (2,3). Hence, a plane is two-dimensional.

    In ordinary, three-dimensional space, we need three coordinates to locate a fly in a room. A fly can be 2 metres away from the left wall, 3 metres away from the back wall, and 1.5 metres up from the floor. Its coordinates are therefore (2,3,1.5).

    This was one of Renรฉ Descartesโ€™s great insights while lying in bed late in the afternoon or so the story goes. Hence, these numbers are called Cartesian coordinates. Descartes was pivotal to the development of what we now call the Cartesian coordinate system.

    The space to which these type of coordinates belong is called Euclidean space as the great Greek mathematician Euclid was the father of Euclidean or classical geometry.

    I think I can safely say that Euclid and Descartes enabled mathematics teachers to torment us with a whole slew of homework in order for us to fully explore the realm of Euclidean space in both two- and three-dimensional Cartesian coordinate systems.

    Figure 3. Home of Descartes in Utrecht, the Netherlands, where he wrote parts of his famous Discours de la Mรฉthode. The house has been demolished. Nowadays, the place looks very different. (Click on the image for a link to the original Instagram post where you can also swipe for the photo of what is looks like today. Opens a new tab.)
    Figure 3. Home of Descartes in Utrecht, the Netherlands, where he wrote parts of his famous Discours de la Mรฉthode. The house has been demolished. Nowadays, the place looks very different. (Click on the image for a link to the original Instagram post where you can also swipe for the photo of what is looks like today. Opens a new tab.)

    Space and time

    In real life, besides a position in ordinary space, you also need to specify when. Getting the coordinates to be inside an office located on the corner of two streets on the 24th floor (that’s the three dimensions of ordinary space right there) just isn’t enough. You also need a time-coordinate. When are you supposed to be there?

    One of my favourite books, Slaughterhouse-Five, or The Children's Crusade: A Duty-Dance with Death by Kurt Vonnegut mentions the Tralfamadorians who โ€˜were friendlyโ€™, and โ€˜could see in four dimensionsโ€™. They also โ€˜pitied Earthlings for being able to see only three.โ€™ They were capable of observing all events at once.
    One of my favourite books, Slaughterhouse-Five, or The Children’s Crusade: A Duty-Dance with Death by Kurt Vonnegut mentions the Tralfamadorians who โ€˜were friendlyโ€™, and โ€˜could see in four dimensionsโ€™. They also โ€˜pitied Earthlings for being able to see only three.โ€™ They were capable of observing all events at once.

    Einstein called the fact that you’re inside an office at a certain time an event. In other words, where, in ordinary, Cartesian coordinates, we talked about some thing being somewhere, Einstein had the insight to now only start talking about events taking place in terms of space and time, space-time โ€“ using space-time coordinates.

    When Einstein introduced the special theory of relativity, the German mathematician Hermann Minkowski realised this theory could also be understood geometrically in a four-dimensional space-time, where time is taken to be the fourth dimension. We now call this space Minkowski space. Note that we’re using the word โ€˜spaceโ€™ in a broader sense: it doesn’t just encompass ordinary spatial dimensions but it now also includes a dimension of time (and, for technical reasons, isn’t Euclidean).

    By the way, another word mathematicians and physicists like to use is manifold. A manifold is a topological object which can take many shapes โ€“ such as a two-dimensional plane, a three-dimensional Euclidean space, four-dimensional Minkowski space or any other space you can mathematically think of.

    In Einstein’s general theory of relativity (gravity), we still work with four-dimensional space-time, except the shape of the space isn’t Minkowskian any more. The shape of the space is warped, curved, and stretched. In the best theory of gravity we have to date, we work on a so-called pseudo-Riemannian manifold, named after the great German mathematician Bernhard Riemann. The dimensions are still all the directions you can take on this manifold, i.e. the minimum amount of coordinates you need to locate an event. However, in this case, they are not necessarily oriented perpendicularly with respect to one another.

    Figure 4. In the film Interstellar (2014), director Christopher Nolan featured an object which had something to do with space and time.
    Figure 4. In the film Interstellar (2014), director Christopher Nolan featured an object which had something to do with space and time. (Click to enlarge.) If you haven’t seen the film and still intend to, do not read this footnote:(beginfootnote)Astronaut Joseph Cooper (Matthew McConaughey) finds himself in this spatial representation of space-time. All four dimensions of particular events in the past, present, and future of a room in his house, chopped up into manageable time chunks, are mapped onto an object (called a Tesseract) inside of a black hole (where the roles of space and time are reversed) through which Cooper can transport himself freely. This enables him to trickle information into the events of his choosing. In the still image above, you see many instances of the same room of his house with his daughter at different positions in time (which is the equivalent of different positions in space for Cooper).(endfootnote). License note.

    Four ordinary space dimensions

    Imagine a Pac-Man living on the surface of a sphere. To them, the world is flat. If they were to travel straight on โ€“ and on and on and on โ€“ eventually, they would be quite surprised to find themselves returning to the point where they started.

    Figure 5. Imagine being as flat as a Pac-Man, travelling on what seems to be a flat surface. You might be surprised to find you'd eventually end up where you started. If you had no knowledge of the three-dimensional concept of a sphere, that is. We do. We know that you'd return because that's what a sphere โ€“ or a circle, for that matter โ€“ does to your path. But what about our Universe? What if we would travel billions and billions of years in a straight line through the Universe? Would we end up where we started? Could our Universe be some kind of hypersphere? (Yes, technically, it's a glome, or an n-sphere, where n=3, and the space it's embedded in is n=1, not an hypersphere. Apologies to the mathematicians and physicists.)
    Figure 5. Imagine being as flat as a Pac-Man, travelling on what seems to be a flat surface. You might be surprised to find you’d eventually end up where you started. If you had no knowledge of the three-dimensional concept of a sphere, that is. We do. We know that you’d return because that’s what a sphere โ€“ or a circle, for that matter โ€“ does to your path. But what about our Universe? What if we would travel billions and billions of years in a straight line through the Universe? Would we end up where we started? Could our Universe be some kind of hypersphere(beginfootnote)Yes, technically, it’s a glome, or an n-sphere, where $n=3,$ and the space it’s embedded in is $n=1$, not a hypersphere. Apologies to the mathematicians and physicists.(endfootnote)?

    We, the three-dimensional beings most of us are, see them as a little surface, a shape, because we can see them โ€˜from aboveโ€™, from the third dimension. We can also see how they’re travelling around the surface of a sphere. They don’t know what a sphere is. They only think of flat surfaces. To us, however, it’s quite logical they would eventually return to their point of origin.

    Okay, so, back to our 3D world. Imagine we travelled in a spaceship, always in a straight line through the Universe. Now imagine, after billions of years, we end up where we started: Earth. What happened? Could our Universe be some kind of sphere, only four-dimensional?

    The cover of the book The Fourth Dimension.
    I can recommend reading The Fourth Dimension: Toward a Geometry of Higher Reality. It became one of my favourite books in the 90s (though it came out in 1984). And there’s of course this book, to which many, such as Carl Sagan and Stephen Hawking, have referred in the past.

    While no experiment has proven the existence of a fourth spatial dimension (let alone five or six etc.), it is a wonderfully entertaining world for the mind to ponder about.

    Just to be absolutely sure: time is not the fourth dimension we’re talking about here. We were talking space โ€“ spatial dimensions. Quite often these two get confused: four-dimensional space-time is three spatial dimensions plus one time-dimension while four-dimensional space is four spatial dimensions without time.

    Abstract spaces

    There’s another way in which dimensions and spaces are used by mathematicians and physicists. Imagine an object having several properties at once: a position (in ordinary space), motion, direction of that motion, temperature, colour. To describe the state of this object, you need more than just four space-time coordinates. Suppose, its space-time coordinates are (0,1,1,1), in other words, it exists at time $t=0$ at position $(x=1; y=1; z=1)$.

    Did we describe the state of the whole object? No, we’re still missing some key properties here. It is in motion, so, it has a speed, say 10 m/s. That speed has a direction โ€“ this is why we say it has a velocity, which is speed and direction. Let’s say its velocity $v = -10 \text{ m/s},$ in other words, it has a speed of $10 \text{ m/s}$ to the left.

    Let’s say its temperature is 273.15 Kelvin, which is 0 โ„ƒ and 32 โ„‰. And its colour is pure white. So, how many numbers do we need to describe the object’s state fully? Exactly, seven numbers (we count โ€˜whiteโ€™ as a number).

    The coordinates (0,1,1,1,-10,273.15,white) are said to live in phase space, an abstract space where the properties of the object form the dimensions of that space. This particular phase space is seven-dimensional. Of course, that’s impossible to imagine, but mathematically, you can work very well with it.

    We gave an unusual example to emphasise that dimensions needn’t be related to spatial and temporal positions. However, usually, phase spaces are indeed used in the context of position and momentum.

    Figure 6. A sample trajectory through phase space is plotted near a so-called Lorenz attractor, a solution to the Lorenz system, which Edward Lorenz developed to model atmospheric convection. The colour of the solution fades from black to blue as time progresses, and the black dot shows a particle moving along the solution in time. The three-dimensional trajectory in phase space is shown from different angles to demonstrate its structure.
    Figure 6. A sample trajectory through phase space is plotted near a so-called Lorenz attractor, a solution to the Lorenz system, which Edward Lorenz developed to model atmospheric convection. The colour of the solution fades from black to blue as time progresses, and the black dot shows a particle moving along the solution in time. The three-dimensional trajectory in phase space is shown from different angles to demonstrate its structure.

    Another example of an abstract space is a so-called vector space where each coordinate does not just occupy a point in that space but that point also has a direction. An example of such a space is the velocity of wind. Each point in that space does not just have a value pertaining to the speed of the air and its location in ordinary space, it has a direction too.

    In the previous post, Complex numbers: an introduction, an entirely new kind of number line was introduced. All the spaces we just mentioned could very well contain complex dimensions. In fact, most of the time, they do. Especially in quantum mechanics. Complex numbers make up abstract complex vector spaces where wave functions thrive. Hilbert space is where it’s at, most of the time.

    The Standard model of quantum physics is based on groups of symmetrical transformations in complex space, called SU(3) $\times$ SU(2) $\times$ U(1). The S stands for special and denotes all possible transformations in complex space except for one particular kind. U(1) refers to a one-dimensional unitary circle group in the complex plane. The numbers indicate the number of dimensions in which these transformations take place. The number of dimensions of the entire system is much higher, though! The dimensionality of the abstract complex space which follows from a symmetry group such as SU(3) is $3^2-1=8.$ As you can see, compared to street corner vernacular, dimensions are very different in scientific context.

    In general, we can say that every manifold is a space. This needn’t pertain to spatial space. The minimum amount of dimensions needed to construct a path to a point on that manifold is the dimensionality of that space.

    There are so many more types of mathematical spaces, they’re too many to mention. Suffice to say, while they have nothing to do with our ordinary space โ€“ our real-world one, which we dwell in โ€“ all these abstract spaces are brilliant mathematical tools enabling us to do predictive calculations pertaining to phenomena taking place in our ordinary, real-world space.

    String theories

    An interesting beast among all of this is string theory. If you accept the premise that an elementary particle such as an electron is actually a spatially one-dimensional string vibrating in specific ways corresponding to the collection of properties of an electron, then more dimensions are automatically needed in order to describe all the particles in this way. Strings need a sufficient amount of freedom, degrees of freedom, to vibrate in unique ways to be able to encompass the entire zoo of elementary particles and their properties.

    Figure 7. The basic building blocks of the entire Universe, according to string theory. Unfortunately, while the theory is mathematically consistent, it cannot yet be (and hasn't been) proven to be correct in this Universe.
    Figure 7. The basic building blocks of the entire Universe, according to string theory. Unfortunately, while the theory is mathematically consistent, it cannot yet be (and hasn’t been) proven to be correct in this Universe.

    In various versions of the string theories, a varying number of dimensions are needed. These dimensions are spatial and invisible. Since we don’t experience these dimensions, it is hypothesised that they are extremely small and curled up. They’re not stretched out like our ordinary three spatial dimensions.

    Or they are so large that to us they don’t affect us in any way noticeable. Just as the curvature of Earth did not affect us when we were little as the Earth is so big compared to our movements.

    Unfortunately, the theory cannot be tested yet. For now it’s purely a mathematical exercise. Although many discoveries have been made in pure mathematics, no experiment has proven string theory to be true (string theory in all its variety, and I’m including superstring theories and M-theory here even though the hierarchy is the other way around). No extra dimensions have been found yet.

    There’s one honourable mention that I’d like to make. It’s the Calabi-Yau manifold, or the Calabi-Yau space. In superstring theory the manifold is hypothesised to encompass six invisible extra dimensions for the theory to work. The manifold is three-complex-dimensional or six-real-dimensional. I like it because it looks cool.

    None of this is proven; we seem to be stuck in this three-dimensional space with one direction of time. And, if you ask me, it’s likely that our three-dimensional space turns out to be a side product of something quantum.

    Figure 8. A Calabi-Yau manifold, named after Eugenio Calabi and Shing-Tung Yau. This is a complex space with complex dimensions. It yields applications in theoretical physics, most notably in superstring theory, where the manifold has six dimensions. Though not experimentally proven to be existing in our world, they do yield fascinating mathematical possibilities and puzzles.
    Figure 8. A Calabi-Yau manifold, named after Eugenio Calabi and Shing-Tung Yau. This is a complex space with complex dimensions. It yields applications in theoretical physics, most notably in superstring theory, where the manifold has six dimensions. Though not experimentally proven to be existing in our world, they do yield fascinating mathematical possibilities and puzzles.

    Spaces and dimensions

    There are so many different spaces with a variety of dimensions that you’d need a whole slew of posts to describe them all properly.

    What can we take away from all of this? Dimensions are not realms. In ordinary space, they are the directions in which objects can freely be transported. That’s three for our world.

    If you model time as a dimension, then we live in a four-dimensional space-time world. Except that you can’t freely move in time as there’s only one direction(beginfootnote)Time is definitely going to be a whole separate set of posts. Can’t wait.(endfootnote).

    Though many had hoped to find extra spatial dimensions, the largest experiment humankind has undertaken, the Large Hadron Collider at CERN, has not found a shred of evidence for them. Instead, it delivered convincing evidence that the current Standard Model of particle physics without extra dimensions is still correct.

    Nevertheless, to describe and predict phenomena in our Universe, it is almost always helpful to model their properties as extra dimensions. This has nothing to do with there actually being extra dimensions โ€“ this is probably where popular and esoteric culture get their inspiration from โ€“ but has everything to do with being able to do calculations in the abstract world of mathematics.

    In a previous post, for example, we assumed imaginary time as an extra dimension to mathematically derive a set of equations in the special theory of relativity. It doesn’t mean imaginary time is an actual extra dimension you can dip appendages or your consciousness into.

    In string theories, actual extra spatial dimensions are required for the theories to work. None of them can be tested as of yet (and none of them have been tested nor proven). It remains to be a beautiful, mathematical construct, but only mathematical.

    In future posts, we will be exploring geometry, pseudo-Riemannian manifolds, symmetry groups, and Hilbert space for loads more bits of maths and physics.

    Licenses

    The featured image in the title and Figures 1 are still images of Marvel Studio’s Doctor Strange (2014) and Figure 4 of Interstellar (2014), all copyrighted films. It is believed that screenshots may be exhibited under the fair use provision of United States copyright law.

    Figure 6. Lorenz attractor animation by Dan Quinn under CC BY-SA 3.0

    Figure 8. Calabi-Yau manifold by Lunch under CC BY-SA 2.5, created in Mathematica

  • Quantum mechanics in ten ideas for people on the move

    Quantum mechanics in ten ideas for people on the move


    The last few posts on quantum mechanics have been quite extensive and at times rather deep for those who are on the move. So, here are ten important ideas about particles and wave functions for when you’re en route in slightly more normal English.


    1: Subatomic particles

    To describe objects in our everyday world, such as rocks, buildings, and cars, Newton’s laws suffice. To describe subatomic, elementary particles, such as electrons, protons, neutrons, and photons, however, there is a whole different type of physics: quantum mechanics.

    2: Wave functions

    The most complete description of an elementary particle is called the wave function. Actually, the word ‘particles’ seems to incorrectly refer to tiny points, balls or spheres or something, which they are absolutely not. They aren’t waves either. ‘Particles’ are wave functions with wave-like properties (emphasis on ‘like’). Upon interaction with other particles and/or measurement, they exhibit particle-like behaviour, however. The wave function contains all physically possible states a โ€˜particleโ€™ can be in at the moment we measure its state. The wave function can be seen as a mathematical description of the probabilities of the states that the particle will snap into as soon as you measure it. It’s often visualised as a โ€˜cloudโ€™ even though that’s not what it actually looks like. It’s just a visual metaphor for a mathematical object that actually lives in complex space as it is complex valued.

    I usually just draw vague spherical thingies.
    I usually just draw vague spherical thingies.

    3: Quantum state

    Once measured, particles show one specific quantum state out of a whole range of possible quantum states prior measurement. Position is the most intuitive to understand example of a quantum state. Momentum is another (momentum is a measure of the amount of motion of a particle). Then there are states such as polarity, spin, and a bunch of others. The wave function encapsulates all these possible states and yields a probability-value for actually measuring a particular state. In other words, even before you measure it, the wave function allows you to calculate the chances of encountering this particular quantum state.

    4: Measurement problem

    As long you don’t measure a particle, and as long as it doesn’t interact with other particles, the particle is not in a specific quantum state yet. Instead, its wave function just describes all these possible states as though they are mathematically added on top of each other. This โ€˜adding of quantum statesโ€™ is what is meant when physicists talk about superposition. The term is from the mathematics of waves and linear algebra in general, not quantum mechanics in particular. While the situation is often portrayed as particles being in multiple states all at once (such as being in two positions at the same time), it’s more accurate to say that the particle does not have a specific state at all. There’s just the wave function with all the probabilities of future quantum states. As soon as you perform a measurement, the particle snaps out of its wave function full of possibilities into a single possibility. In other words, what you see is not what it was. What you observe is just a sliver of its total prior existence. How this happens, nobody knows. It’s called the measurement problem. There’s a Nobel Prize waiting for you.

    This extraordinary experiment yielded a photo of the closest approximation of the wave function of an electron in a hydrogen atom we have to date. It was made by the Polish physicist Aneta Sylwia Stodolna et al. (Source: Stodolna AS et al. (2013) โ€œHydrogen Atoms Under Magnification: Direct Observation of the Nodal Structure of Stark States,โ€ Physical review letters, 110(21), pp. 213001โ€“213001.)
    This extraordinary experiment yielded a photo of the closest approximation of the wave function of an electron in a hydrogen atom we have to date. It was made by the Polish physicist Aneta Sylwia Stodolna et al. (Source: Stodolna AS et al. (2013) โ€œHydrogen Atoms Under Magnification: Direct Observation of the Nodal Structure of Stark States,โ€ Physical review letters, 110(21), pp. 213001โ€“213001.)

    5: Schrรถdinger equation

    Wave functions obey the Schrรถdinger equation. You could say that what Newton’s second law is for objects in our everyday world, is what the Schrรถdinger equation is for the subatomic world. It gives us the ability to predict how the wave function evolves in time. This is a completely classical equation; it is 100% deterministic. Where the wave function captures a range of probabilities, the Schrรถdinger equation tells us how this range of probabilities changes over time perfectly predictably so. In other words, it doesn’t predict the exact state of a particle once measured, but it does accurately predict the probability-value of an exact state once measured at any given time.

    6: Uncertainty

    There is a fundamental informational trade-off between certain possible states such as between position and momentum, energy and time, and time and frequency. The origin for this does not lie in quantum mechanics. It’s due to the way they are related to each other. Mathematically, these variables are called Fourier transform pairs or conjugate variables. To calculate one from the other, you have to execute a mathematical procedure called a Fourier transform. The trade-off is that Fourier transforming a variable whose range of possible values is smaller leads to the other variable having a larger range of possible values. And if a range of possible values becomes larger, then the exact outcome of measurement is less certain (the probability of a specific state after measurement becomes more uncertain). Heisenberg showed that this uncertainty principle also applies to the wave function in quantum mechanics, hence, there the principle is called Heisenberg’s uncertainty principle.

    Fourier showed that if a sound is fairly well-defined in time (bottom), it has to be comprised of multiple frequenties (illustrated as multiple waves at multiple frequencies). That's the fundamental uncertainty principle with waves.
    Fourier showed that if a sound is fairly well-defined in time (bottom), it has to be comprised of multiple frequenties (illustrated as multiple waves at multiple frequencies). That’s the fundamental uncertainty principle with waves.

    7: Certainty

    That same principle predicts that, while very valid at the scale of subatomic particles, this uncertainty becomes utterly meaningless at our large-scale world of everyday objects. A bowling ball whose range of possible positions is very limited (locked in a very tight enclosure with little to no leeway), will never portray any uncertainty values pertaining to its motion (momentum), for instance. By Heisenberg’s uncertainty principle, upon measurement, it might show to have the speed of $3.283 \times 10^{-35} \text{ m/s}.$ This means that after 965.9 billion years it will have travelled the distance of the width of a proton. So, no, uncertainty effects play no role in our everyday world, unless you are doing experiments with a running time of seventy times the age of our current Universe. In that case, you will have to deal with the uncertainty of the width of a proton(beginfootnote)When people state or think that everyday objects (our bodies, brains, tennis balls, animals) can exhibit quantum effects such as being at multiple places at the same time, I suspect this is because they have no well-defined idea how small subatomic particles really are and no inkling as to how large the everyday world is in those terms. Also, they didn’t do the calculations.(endfootnote). We do note that extraordinarily sensitive larger-scale equipment such as the the mirrors at the LIGO and Virgo experiments are capable of measuring quantum effects, however, this isn’t really unexpected nor is it the same as saying a human body is in a quantum superposition. Measuring quantum effects is one thing, brains supposedly being in two places on Earth (โ€˜based on principles from quantum mechanicsโ€™) is a whole other thing.

    Missing the pins has nothing to do with practical nor theoretical quantum effects. You're just not that good.
    Missing the pins has nothing to do with practical nor theoretical quantum effects. You’re just not that good.

    8: Quantum entanglement

    When two or more particles can only be described by one wave function โ€“ not as separate wave functions โ€“ those particles are said to be quantum entangled, either partly or completely. A measurement performed on one particle immediately determines the measurement outcome on the other entangled particle, irrespective of the spatial distance between them. This is why this phenomenon is said to be non-local. How this happens, is unknown. This effect dissipates to zero when entangled particles interact with yet other particles. At the large scale of our everyday world, the number of particles inside an object to be interacted with is so great, quantum entanglement completely fades away. In very special conditions, however, such as in our labs, entanglement can be sustained for quite some time.

    This isn't what quantum entanglement looks like. It's just a picture.
    This isn’t what quantum entanglement looks like. It’s just a picture.

    9: Quantum Field Theory

    Over the years, the mathematical and physical theory of (quantum) wave mechanics has been extended to describe quantum fields as the fundamental building blocks of our Universe. The Universe is made of quantum fields. The most complete description of fields are wave functions. This is called Quantum Field Theory (QFT). The most successful version of QFT is called the Standard Model of quantum physics. โ€˜Particlesโ€™ are here some kind of disturbance in their field: an electron is a disturbance in the electron field. The particle’s description is here part of the wave function of its entire field. The challenge is now to extend this quantum field theory into its next form, encapsulating something called quantum gravity. What we don’t know yet, for example, is how to have space and time in extreme regions such as black holes, naturally appear out of a quantum theory.

    My very sketchy way of showing quantum fields. Proton field is not really a thing. It's just a shortcut for several quark fields. Besides, fields aren't two-dimensional, they're obviously three-dimensional.
    My very sketchy way of showing quantum fields. Proton field is not really a thing. It’s just a shortcut for several quark fields. Besides, fields aren’t two-dimensional, they’re obviously three-dimensional.

    10: Applications

    While the famous physicist and Nobel Prize winner Richard Feynman is known for having said, โ€˜I think I can safely say that nobody understands quantum mechanicsโ€™, this is sometimes incorrectly taken to be a reason to state that, therefore, physicists don’t know what they’re talking about. Feynman alluded to the fact that there is much we don’t know about the foundations of quantum mechanics. There is still much employment in solving hard problems such as quantum gravity, the measurement problem, the strong CP problem, the interpretation of quantum mechanics, non-locality, and so forth.

    On the other hand, we now have WiFi, internet, touchscreens, lasers, MRI scanners, LEDs, flash memory, solid state disks, the old crunchy hard disks, transistors, and CPUs or integrated chips (ICs) in general.

    I think I can safely say that the fact that you’ve plucked this article out of the air to have it displayed on your (touch)screen is at least an indication of the level at which โ€˜nobody understands quantum mechanicsโ€™.

    Nevertheless, we’re far from done. There is still much to discover in this Universe with a bit of maths and physics.

  • Complex numbers: an introduction

    Complex numbers: an introduction


    Complex numbers have fascinated me since high school. Usually, it’s where we are taught about natural numbers, integers, rational, irrational, and real numbers but never about complex numbers. This post is for those who might be interested in an easy introduction into the realm, or rather, plane of complex numbers. And they’re not without practical significance either: no electronic device such as the one you’re using to read this post could have been built without physicists, electrical engineers, and computer scientists knowing anything about the gift of complex numbers from sixteenth century mathematicians.

    Blown away

    โ€˜There are such things as negative numbersโ€™, explained my father to me when I must have been about six or seven years old since I was a second-year pupil in primary school. He explained the notion of negative possession when owing a certain number of marbles to someone which was greater than the number of marbles you physically carry with you. As this was one of those I-still-remember-where-I-was-when moments, like it was yesterday, I remember sitting on the floor besides the coffee table in the living room of our terraced house in the town of Emmeloord, which had been reclaimed just forty-three years earlier from the IJsselmeer, a lake formerly part of the North Sea.

    I clearly remember feeling exactly the same when he had told me earlier our planet wasn’t flat and when my Mum told me in the car yet a few months earlier, that we were living on the sea floor. The cap of my mind was blown away, yet again. It took a while before I managed to fold my slow and wet brain lobes around the notion that negative numbers existed, even though you couldn’t see them in the real world like you could โ€˜seeโ€™ regular numbers such as in lengths or the number of marbles(beginfootnote)Inexplicably, I had never considered the fact that temperature could get below 0 โ„ƒ, which it still did, back then in The Netherlands. We used to enjoy an outdoor activity called ice skating, on frozen lakes, ponds, rivers, and ditches.(endfootnote).

    I hastened to tell my primary school teacher excitedly about negative numbers. She just nodded and then told me to proceed with doing my homework on boring regular arithmetic. She had a point as I wasn’t very good at it.

    Fast forward to when I must have been about fifteen or sixteen when I read about complex numbers in a popular textbook about quantum mechanics. The fact that they were called โ€˜complexโ€™ may have triggered my curiosity as I assumed that term pertained to it being very difficult, but mostly because, apparently, so-called imaginary numbers are a thing! I had that exact same feeling again. The cap of my mind had melted. The whole notion seemed to radiate some kind of magical power. What sorcery was this? Could this be a doorway to extra dimensions?

    The next day, I told my mathematics teacher, Mr Es โ€“ Es is not his actual name but it was his two-letter code in our high school timetable. I’ve always found it appropriate Es is also the symbol for the element Einsteinium in the periodic system. As his first name happened to be the same, my friends and I used to joke that we were on our way to the lessons of Albert Einstein.

    Mr Es did what every good teacher does when a student tells you something they get enthusiastic about: he encouraged it โ€“ in his case by lending me his old textbook from when he was a first-year mathematics student in Amsterdam. It was an introductory text about complex numbers at the level of undergraduate mathematics.

    The very textbook. (Click to enlarge.)

    I’m ashamed to say I kept it. It was one of those instances where, after the nth time of moving house, I realised, oh my god, I still have this!? It’s also true that I treasured it. It carries a special meaning to me. It signifies how, at least once in my lifetime, I felt acknowledged in what stirred me deeply at the time. A thing I couldn’t really share with friends or anyone close in general, I suddenly shared with someone very clever whose name was denoted by the symbol for Einsteinium.

    Thanks to the miracle of internet, we got back in touch, about twenty-five years later. I confessed I had always kept it and apologised. He had indeed wondered where it had been as he once wanted to show it to someone else. But I could keep it as he was cleaning out the attic anyway. And he was glad it had done something for me as he learnt about my current engagements in a bit of maths and physics.

    I felt guilty. I still do. Someone else could have enjoyed it just as much as I have. And now I have prevented that from happening through his book. So, whoever you are, my sincerest apologies.

    I hope, one day, I will be able to ignite sparks of joy for the beautiful mathematics of complex analysis to many others. I also hope you might experience at least a fraction of the amazement I felt and that the newly gained insight on the concept of โ€˜numbersโ€™ might turn out to be beyond what you were able to imagine so far. So, let this be a beginning.

    Number sets

    A game of hopscotch drawn on the pavement with numbers on the tiles

    We all know and love (or hate, depending) the natural numbers: the whole numbers we count things with. 1, 2, 3, etc. Some mathematicians will want to include the number 0 while others don’t. In any case, this mathematical set of numbers is called the natural numbers and is denoted by the symbol $\mathbb{N}$.

    Then my father told me about the negative numbers, such as -1, -2, -3, etc. If you include the natural numbers and add to that these negative numbers, and add the number 0 to it (if you hadn’t already), then the result is an entirely new set of numbers called the integers, denoted by the symbol $\mathbb{Z}$.

    To denote that the set $\mathbb{N}$ is part of the larger set $\mathbb{Z}$, people use this symbol for subset, $\subset$. They will write $\mathbb{N}\subset\mathbb{Z}$, the natural numbers are a subset of the integers.

    Of course, there’s the ratio’s. The fractions. Between 1 and 2, there’s 1.5. So, in fraction-notation, that’s $\frac{3}{2}$. They’re obviously not whole numbers. They’re rational numbers because they can be represented by a ratio of integers. This number set is symbolised by $\mathbb{Q}$. We now have $$\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}.$$

    It is interesting to note that, therefore, by this expression of subsets of subsets, even numbers such as 9 are rational numbers. On the surface, it’s not a fraction. Below the surface, however, it can be expressed as a ratio of integers: $9=\frac{9}{1}=\frac{18}{2}=\frac{36}{4}$, for example (and infinitely more).

    But wait, there’s more. Fractions such as 1.5 and 3.2 are finite. What if the decimals don’t end? What if you can’t write a particular kind of numbers as ratios, such as with the number $\pi$ or $\sqrt{2}$? These numbers are called the irrational numbers. They are all the numbers which aren’t rational. There’s no symbol for that(beginfootnote)Often, mathematicians circumvent the lack of a symbol by writing something like โ€‹โ€‹โ€‹$\mathbb{R} \backslash \mathbb{Q}.$(endfootnote).

    Instead, there’s a symbol for all the natural numbers, the integers, the rational numbers, and the irrational numbers altogether(beginfootnote)Yes, indeed, my dear fellow mathematician, you thought correctly, I am skipping transcendental numbers here (and algebraic numbers, for that matter). As all transcendental numbers are irrational numbers but not all irrational numbers are transcendental, I decided it over-complicated things in what was supposed to be an introductory text on complex enough numbers anyway.(endfootnote). They’re called the real numbers and this set is denoted by $\mathbb{R}$. This is the set we’re all used to working with. We now have $$\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}.$$

    The set of real numbers $\mathbb{R}$ contains all the numbers. Or does it?

    A diagram of all the number sets in the shape of ellipses. The ellipse of R containing the ellipse of Q containing the ellipse of Z containing the ellipse of N.

    The secret of del Ferro, del Fiore, Tartaglia, and Cardano

    Well, you guessed it. Here they come, the complex numbers. Let’s do just a tiny bit of maths. Remember what the quadratic of a number was? And what a square root was? What is the square root of 64, in other words, $\sqrt{64}$? Yes, that’s 8. Because 8 times 8, or 8 squared, or $8^2$ equals 64.

    Okay, suppose $x^2 = 64$, what is $x$ then? Well, you do exactly the same thing, you un-square $x$ by taking its square root. And you have to do the same with the number after the equal sign. So, $\sqrt{x^2} = \sqrt{64}$, in other words, $x = 8$.

    Tartaglia

    Maybe you remember this comes in handy when calculating the lengths of the edges of your piece of land. Suppose, the surface area of your square piece of land is 64 square kilometre (or square miles). What is the length of an edge of that land? That’s 8 kilometre (or miles).

    All these calculations take place in the realm of $\mathbb{R}^+$, the positive part of all real numbers. Note that no surface area of a piece of land can be negative. In other words, a surface area of -64 square metres is nonsensical. Also, the square root of -64 has no solution. It’s not -8, because -8 times -8, or $(-8)^2$ is simply 64 again, because a negative number times a negative numbers equals a positive number as we proved in an earlier post.

    Cardano

    Sometime in the sixteenth century, somewhere in Italy, Scipione del Ferro, professor of the University of Bologna, solved a slightly different kind of equation. It was a so-called cubic equation. Where we basically found the solution to a quadratic equation such as $x^2 = 64$ from the top of our heads, he found solutions for a cubic equation such as $x^3 + x^2 + 6x + 3 = 0.$ Del Ferro was known for not wanting to publish any of his proofs and solutions. He kept a secret notebook and that was it.

    On his death bed, however, he told his pupil Antonio Maria del Fiore the secret to solving it. Del Fiore went on to challenge Niccolรฒ Fontana Tartaglia, a mathematician residing in Venice at the time. Tartaglia had actually solved it himself before and trusted the formula to Gerolamo Cardano, the then Milan-based polymath and genius. Tartaglia messaged the solution in the form of a poem (no less!) but didn’t entrust the proof to him.

    Of course, Cardano was able to reconstruct the proof anyway. As he learnt that del Ferro had also found the solution, he then proceeded to publish it all in his Ars Magna from 1545, much to the chagrin of Tartaglia.

    So, what was the secret so many large minds had been secretive about? A new type of number.

    imaginary

    Let’s take a simpler example. Suppose, we have the following simplistic quadratic equation: $x^2 – 4 = 0$. To solve it, we ‘move’ the 4 to the other side of the equal sign, by adding 4 to both sides: $x^2 – 4 + 4 = 0 + 4$, which simply becomes $x^2 = 4$. If you apply the square root to both sides, you get $\sqrt{x^2} = \sqrt{4}$. The solution to this equation is thus $x=2$ or $x=-2$ (because $-2\times -2 = 4$ too).

    Good. Basically, the mathematicians of the sixteenth century opined that they should be able to solve a variation of this equation as well: $x^2 + 4 = 0$. Let’s bring the 4 again to the other side of the equal sign by subtracting 4 on both sides: $x^2 + 4 – 4 = 0 – 4$, which becomes $x^2 = -4$. Now, again, the question is, what is $x$?

    Let’s try and apply the square root to both sides again: $\sqrt{x^2} = \sqrt{-4}$. Halt. Stop. What is the square root of -4? What is the square root of a negative number?

    We have the same situation where we are to apply the square root of a negative surface area. The answer isn’t -2, because $-2\times -2 = 4$, not -4. What then?

    Before del Ferro, Tartaglia, and Cardano, people would have said that there simply is no solution. Thanks to them, however, we can solve it. The answer lies in the following definition: $$i^2=-1.$$

    This seemingly simple act enables us to solve $x^2=-4$. We can then write $x = 2i$ or $x = -2i$.

    Let’s take our first solution, $x = 2i$. If we square this, we get $x^2 = (2i)^2$, which we can also write as $x^2 = 2^2i^2$. Now, since $i^2 = -1$, we can substitute that to get $x^2 = 2^2(-1)$, which is, of course, $x^2 = -4$. Ecco!

    The same goes for the other solution, $x = -2i$. If we square this, we get $x^2 = (-2i)^2$, which we can write as $x^2 = (-2)^2i^2 = 4i^2 = 4(-1) = -4$. Ecco!

    So, you may ask, what devilish entity is this $i^2=-1$? The letter $i$ stands for โ€˜imaginaryโ€™ and so, $i$ is a so-called imaginary number.

    Now, because $i^2=-1$, you can also write(beginfootnote)Although, I actually prefer to use $i^2=-1$ over $i=\sqrt{-1}$ even though the latter has been mentioned in many school books. However, I believe it might lead to confusion. Since we have the rule that $\sqrt{a}\sqrt{b}=\sqrt{ab}$ where $a$ and $b$ are positive real numbers, you might try to apply this rule to negative real numbers, such as when $a=b=-1$. You would then get the incorrect statement $\sqrt{-1}\sqrt{-1} = \sqrt{(-1)(-1)} = \sqrt{1} = 1$, which is wrong as it should be equal to -1. That’s why I try to avoid using $i = \sqrt{-1}$ where I can.(endfootnote) that $i = \sqrt{-1}$. And that’s the crazy part: how can you calculate the square root of a negative number? How can you calculate the square root of a negative surface area? The answer is, you can’t. Not in the realm of the real numbers $\mathbb{R}$, that is. However, we’re not in Kansas anymore, Dorothy. We’re in a new land called the complex numbers. Bye $\mathbb{R}$, and welcome to $\mathbb{C}$.

    Here are some examples of complex numbers: $2i$, $\frac{2}{3}i$, $i\sqrt{2}$, $i \pi$, $-0.25i$. What’s more, you can add these imaginary numbers to a real number such as 3, like so: $3 + 2i$ or $3 + \frac{2}{3}i$ etc. These sums are their own answer. They are complex numbers.

    A complex number $z$ is of the form $z = a + bi$, where $a$ and $b$ are real numbers and $i^2 = -1$. The first real number, $a$, is called the real part of $z$. The last real number, $b$, is called the imaginary part of $z$. The set of all complex numbers is denoted by $\mathbb{C}$.

    And so, we now have

    $$\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}\subset\mathbb{C}.$$

    Note that every real number is a complex number but not every complex number is a real number. That is what one thing being a subset of another thing means. For instance, the real number 9 is a complex number where $b=0$. In other words, the real number 9 can be written as the complex number $9 + 0i$, which is simply 9, which thus happens to be a real number too.

    But $z = 3 + 2i$ is not a real number, because it has an imaginary part which is not equal to zero. So, $z$ is now exclusively a complex number.

    A diagram of all the number sets in the shape of ellipses. The ellipse of C containing the ellipse of R containing the ellipse of Q containing the ellipse of Z containing the ellipse of N.

    Complex plane

    Graphically, all the real numbers of $\mathbb{R}$ can be thought of as a point on the number line.

    A diagram depicting the real number line. Every point on this line represents a real number, such 0, 1, 2, 3 and the square root of 2, pi, and e.

    So, where do complex numbers reside?

    Owing to people such as Wallis, Wessel, Argand, Buรฉe, Mourey, Warren, Franรงais, Bellavitis, Gauss, and Euler[1], the idea to extend the real number line with an imaginary number line perpendicular to the real number line came to fruition. What you get is the so-called complex (geometric) plane, sometimes called the $z$-plane, Gauss plane or Argand plane.

    So, a complex number such as $z = 3 + 2i$, โ€˜containsโ€™ the real number $3$ along the real axis, and the imaginary part, along the imaginary axis, sits at $2i$. A complex number is therefore always represented by a point in a two-dimensional space. Note that all the numbers from all the subset of complex numbers, i.e. $\mathbb{R}$ all the way down to $\mathbb{N}$, can also be represented by a point in this same two-dimensional complex space โ€“ it’s just that they all reside on the real axis.

    As you can -heh- imagine, doing calculations with complex numbers has become an exercise of geometry now! In fact, one of the most beautiful equations in mathematics (at least to my taste) pertains to trigonometry in the complex plane; it’s called Euler’s Formula.

    A diagram representing the complex plane. Perpendicular to the real number line is now a so-called imaginary axis with numbers such as i, 2i, 3i, pi-i, i square root of 2, etc. A complex number is now a point in on that surface.

    Not so imaginary

    It’s unfortunate that this number $i$ and any real number multiplication of it are called imaginary numbers. It was the renowned French philosopher and mathematician Renรฉ Descartes who coined the term imaginary numbers because he considered them to be illusory. In fact, even Cardano had described them as โ€˜some recondite third kind of thingโ€™[2].

    It’s unfortunate because โ€˜imaginaryโ€™ leads to semantic ambiguity. I get it: you would never see something like $\sqrt{-1}$ in the real world. But neither would you see $\sqrt{2}$ out in the wild, for that matter. And yet, it’s the exact length of the hypotenuse of a particular right triangle, which a skilled DIY person could make while you’re waiting. To me, โ€˜realโ€™ numbers such as $\pi = 3.1415926535897 \dots$ without ever ending are as real as โ€˜imaginaryโ€™ numbers are (and vice versa). Circles are a real thing and $\pi$ can be used to do calculations on them. Well, with imaginary numbers you can do calculations on them just as well.

    Complex numbers are used in a variety of sciences. In Einstein’s relativity, which makes GPS navigation possible, you could make use of so-called imaginary time. This sounds like a concept straight from a science-fiction novel, however, imaginary time is a well-defined concept. In fact, in a previous post, we used this to derive the central set of equations in relativity, called the Lorentz transformations. See how the word โ€˜imaginaryโ€™ might invoke unwanted ambiguity?

    To make quantum mechanics work โ€“ the most successful theory to date โ€“ complex numbers are all over the place. Without them, the computer, mobile phone, tablet, TV, VCR, even your modern fridge โ€“ they wouldn’t have worked as no engineer would have been able to produce integrated circuits. The wave function is a complex function living in a complex separable Hilbert space, taking on complex probability amplitudes, evolving according to the Schrรถdinger equation, which itself is a complex equation.

    In mathematics, one of the better-known areas of research where complex numbers play a central role is the study of complex dynamical systems. The featured image above is a detail of the famous Mandelbrot set. It’s a special collection of complex numbers, the projection of which you see plotted colourfully in the complex plane. The study of (complex) fractals also informs all kinds of patterns in nature and growth, even weather forecasts, and climate science โ€“ they’re all informed by complex-dynamical areas of mathematical interest. Also, we’ve used them in a previous post, calculating whether a lab centrifuge with $n$ available spots can be balanced out by a $k$ number of test tubes.

    A fun application of complex numbers is computer games. To calculate rotations in three-dimensional space, computer scientists make use of quaternions, which are an extension of the complex plane. A quaternion is an expression of the form $a + bi + cj + dk$, where $a,b,c,d$ are any old real numbers, and $i^2=j^2=k^2=-1$. However, this is perhaps an interesting subject for another bit of maths and physics.

    [1] Cooke, R. (2005) The history of mathematics : a brief course. 2nd edn. New York, N.Y.: Wiley.

    [2] Open University (2014) Essential mathematics 1. Milton Keynes: Open University.

    Images

    Featured image: Mandelbrot set – Step 6 of a zoom sequence by Wolfgang Beyer under CC BY-NC-SA 2.0; adapted to fit layout.

    Hopscotch Game by ncassullo.

    Niccolรฒ Fontana Tartaglia. Rijksmuseum, Dutch National Museum. Public domain.

    Girolamo Cardano. Wellcome Images under CC BY 4.0.

  • Heisenberg’s uncertainty principle

    Heisenberg’s uncertainty principle


    It’s perhaps not as famous as Einstein’s formula but in this day and age many people may still have heard at least once of the phrase โ€˜Heisenberg’s uncertainty principleโ€™. It plays an important role in quantum mechanics. You may have heard that every time you observe or measure matter, due to the crudeness or inherent inaccuracy of the measurement device, you will inevitably disturb your own observation. This would then preclude you from gaining accurate knowledge with satisfying certainty. In fact, in general, Heisenberg’s uncertainty principle states that nothing can be certain. At the risk of sounding vague and vanilla, all of these statements are completely and utterly wrong. Let’s look at what it really says, shall we?

    Figure 1. Werner Heisenberg in Gรถttingen in 1924.

    Fourier transform pairs

    Trade-offs. Who doesn’t hate them? Remember when your parents told you that you could have this but then not have that or maybe just a bit of this but then less or fewer of that? Unsurprisingly, at least three famous philosophers have written a few words on this, each in their own way lamenting on the existence of trade-offs and how to deal with them. One chose to become all rebellious about it and wrote: โ€˜I want it all, I want it all, and I want it now!โ€™ (May, 1988). The other two, however, chose to be more pragmatic about it as they postulated that โ€˜you can’t always get what you wantโ€™ (Jagger & Richards, 1968). Obviously, they knew that, sometimes, life brings you Fourier transform pairs. The more well-known example is of course Heisenberg’s uncertainty principle.

    If you limit a particle’s range of possible positions in space $(\Delta x)$, you increase its range of possible momenta(beginfootnote)Momentum is the product of mass $m$ and velocity $v,$ so $p=mv.$ It’s a measure for the amount of motion of an object.(endfootnote) along the $x$-direction $(\Delta p_x),$ and vice versa.

    This is formalised as follows:

    $$\Delta x \Delta p_x \geq \frac{\hbar}{2}.$$

    Just to be absolutely clear: the delta-symbol $\Delta$ is a range of a certain quantity. Usually, a $\Delta$ is defined as the difference between two values. Suppose, you measure point $A$ of your garden fence to be $0.1$ metre away from your wall and point $B$ to be $0.7$ metre away from your wall, then the $\Delta$ of the distances, i.e. the length between points $A$ and $B,$ is $0.7-0.1=0.6$ metre.

    In Heisenberg’s principle, it is stated that the product of the range of possible positions $\Delta x$ and the range of possible momenta $\Delta p_x$ is greater than or equal to some number. Mind you, it’s a tiny number. The symbol $\hbar$ stands for the Planck constant divided by $2 \pi,$ and the result gets cut in half yet again.

    This means that whenever one is getting bigger, $\Delta p_x$ for instance, the other is getting smaller, which is then $\Delta x.$ And vice versa.

    Click here if youโ€™d like to do a bit of maths. It’s very easy.

    Just to get an intuitive insight in this relation, suppose $\frac{\hbar}{2}=1,$ and so, suppose, $\Delta x \Delta p_x = 1.$ Furthermore, suppose $\Delta x = 0.5.$ What value does $\Delta p_x$ has to be to satisfy this equation? Exactly, $\Delta p_x$ has to be $2,$ because $0.5 \times 2 = 1,$ or else the equation is false.

    Now, lets make $\Delta x$ smaller. In other words, we’re going to try to pinpoint the location with much more precision. So, let’s say, $\Delta x = 0.001.$ What value does $\Delta p_x$ has to become to satisfy this equation? You guessed right, $\Delta p_x$ has to become even larger: $\Delta p_x = 1000,$ because $0.001 \times 1000 = 1.$ If you were to reverse the situation โ€“ decreasing the size of $\Delta p_x$ โ€“ then, in turn, $\Delta x$ would have to become larger.

    In reality, $\frac{\hbar}{2}$ is much smaller than 1. It is, in fact, about $5.273 \times 10^{-35} \text{J/s}.$ That’s thirty-four zeros behind the decimal point and then ending in 5273. It’s incredibly small. Don’t worry about this. We’ll get back to that later.

    Hopefully, now you see the relation between $\Delta x$ and $\Delta p_x$ as put forward by Heisenberg’s formulation. They complement each other. Whenever a range of possible values becomes larger, in other words, the $\Delta$ or range of value-options is larger โ€“ its actual value becomes more uncertain, hence the use of the word โ€˜uncertaintyโ€™ in Heisenberg’s uncertainty principle(beginfootnote)In fact, it’s statistics. The $\Delta$-sign could just as well be a $\delta$-sign, so $\delta x \delta p_x \geq \frac{\hbar}{2},$ which signifies its statistical character more accurately. After all, the wave function is about probabilities.(endfootnote).


    But why is this? While this principle plays a central role in quantum mechanics, it’s actually not fundamentally a quantum-mechanical law. This principle exists more generally in many instances in physics, and, even more generally, in mathematics.

    In mathematics, the variables position and momentum are said to be a Fourier transform pair. Put in yet other mathematical jargon, position and momentum are said to be conjugate variables.

    Sound

    A well-known, non-quantum-mechanical example of the uncertainty principle is determining the pitch of a sound. How โ€˜highโ€™ a note is, depends on the frequency.

    The most familiar way we depict sound waves is a simple sine wave. It represents the simplest of sounds possible. Also, it’s the most boring of sounds possible.

    The $x$-axis represents time. The $y$-axis represents the amplitude of the sound or the loudness, the intensity of it. As you can see, the sound wave repeats itself over time; the pattern is cyclic. One whole cycle is when the plot has completed going up, going down, going further down, and going up again. The time it takes to complete one cycle is designated by the symbol $T,$ called the period(beginfootnote)It is also possible to measure the time-distance between two peaks or two troughs.(endfootnote). So, this particular sound wave is said to be periodic.

    Figure 3. A time-amplitude plot of a boring old sinusoidal sound wave. (Click to enlarge.)

    The shorter the period ยญโ€“ the quicker the cycles are โ€“ the higher the tone. Another way of saying, is that the higher the frequency, the higher the tone. The mathematical relationship between period $T$ and frequency $f$ is the following expression:

    $$f = \frac{1}{T}.$$

    If the period gets shorter, i.e. the value of $T$ becomes smaller, then the value of $f$ becomes larger, which means higher, which means a higher tone.

    Seeing as the time period $T = 2 \pi$ seconds, the frequency diagram looks like a spike at $\frac{1}{2 \pi}$ Hz. In this frequency-diagram, the $x$-axis is the frequency and the $y$-axis is still the amplitude.

    Figure 4. A frequency-amplitude plot of the sound wave of Figure 3. It shows the exact frequency at which that sound wave exists.

    So, there are now two ways in which we can describe the sound wave: either by frequency (Figure 4) or by change over time (Figure 3).

    Notice that the sound wave plotted as a function of time (Figure 3) has no beginning nor end. For all we know, that plot could just go on forever, to an infinite amount of time, in both directions. Suppose, we would ask the question at what time exactly does the sound exist? The answer is: always. There is no particular, specific time at which it exists.

    In other words, we could write that $\Delta t = \infty.$

    Notice, however, that the frequency plot looks very finite: just one stroke. One well-defined, finite stroke. If we were to ask the question what frequency exactly does the sound have? The answer is: there is a particular, specific, exact frequency at which it exists and it is $\frac{1}{2 \pi}$ Hz โ€‹
    $( \approx 0.16).$

    Fourier analysis

    In reality, no sound is going to be infinitely long. Pluck a guitar string and it will fade out as the energy dissipates slowly. Also, at some point it started ยญโ€“ meaning, before that, it didn’t exist. In other words, in reality, a sound wave usually exists in a finite range of time.

    Let’s limit our sound wave to a range in time, so it looks more like the sound of a โ€˜blipโ€™ and less like an infinite tone of boredom. Again, the $x$-axis represents time and the $y$-axis represents the amplitude.

    Figure 5. A time-amplitude plot of a so-called wavelet, a short sound burst. Contrary to the sound wave in Figure 3, it’s not infinitely long. It’s now also more difficult to assess its frequency.

    As you can see, the sound now exists in a more defined range of time โ€“ roughly 1.5 seconds. In other words, $\Delta t \approx 1.5$ seconds. That’s a whole lot smaller than the old $\Delta t = \infty.$

    Now, we ask ourselves, what is its frequency? The difficulty now is that it’s hard to pinpoint an exact period $T$. The evolution of the plot is quite different from our infinitely long sine wave. Yes, we can identify kind of those cycles we’re looking for, however, no cycle has the same shape, so, technically, we’re dealing with multiple cycles at once. And guess what, its frequency-amplitude plot looks like this.

    Figure 6. The frequency-amplitude plot of the wavelet in Figure 5. It’s far from being a specific, exact frequency. At varying degrees, it’s actually a few frequencies at the same time.

    As you can see, it has become difficult to pinpoint the exact frequency of our wavelet. It exists at a variety of frequencies and amplitudes.

    So, while the โ€˜time windowโ€™ of the sound wave has become more exact, the frequency has now become โ€˜less certainโ€™.

    The brilliant mathematician Joseph Fourier discovered that a wavelet such as in Figure 5 can actually be constructed by adding many infinite waves at many frequencies. Put differently, Fourier analysis shows that our wavelet is the culmination of a superposition of many waves at many frequencies.

    Figure 7. The wavelet at the bottom is constructed by many infinite waves at many different frequencies superposed onto each other. This automatically means that the wavelet’s exact frequency is fundamentally harder to determine than the frequency of the sound wave in Figure 3.

    Now you see why the frequency-amplitude plot has changed from a very specific value in Figure 4 to the wider set of frequencies in Figure 6. In the latter case, the wavelet โ€˜containsโ€™ multiple waves at multiple frequencies, so when you Fourier transform its time-amplitude plot to its frequency-amplitude plot, the frequency has become โ€˜uncertainโ€™.

    The relation between time $\Delta t$ and frequency $\Delta f$ in ordinary classical physics is fundamentally complementary. No quantum mechanics needed.

    In mathematical jargon, time and frequency are so-called Fourier transform pairs or conjugate variables.

    The term โ€˜Uncertainty principleโ€™ pertains to the general phenomenon that Fourier transforms (such as between time and frequency) entail a fundamental, mathematical trade-off between types of information carried by the two transformed variables. Heisenberg then showed that this principle also holds in quantum mechanics. And so, the uncertainty principle in quantum mechanics is called Heisenberg’s uncertainty principle.

    The De Broglie relation

    Time to go back to quantum mechanics. Remember that a particle’s best description is a wave function? A wave function is the mathematical expression of a particle containing all possible states it can assume once we measure it.

    Instead of a time-amplitude plot, let’s represent a particle by a space-amplitude plot. To make it a little bit easier, let’s take the wave function of a particle of which the amplitude only varies along one dimension of space, $x.$

    Here is a representation of a particle’s wave function along one dimension of space (along a โ€˜straight lineโ€™). The $x$-axis represents a position in space. The $y$-axis represents the amplitude of the wave function (which is proportional to the probability of finding the particle in that particular position $x$).

    Figure 8. A representation of a wave function of a free particle. Note that this is not what it actually looks like. For one, an actual wave function exists in complex space, which we didn’t plot here. The goal is to illustrate, not to map accurately. Also note that the free particle has no specific position yet as it’s a free particle!

    It was the eminent French physicist Louis de Broglie(beginfootnote)Many physicists have tried and mispronounced his last name. It should sound like โ€˜broyโ€™ where the r is produced at the back of the throat, like the French r โ€“ a โ€˜dryโ€™ kind of r. In this interview with him, you can hear the French presenter pronouncing his name (just after 0:16 seconds). It’s not โ€˜brog-lyโ€™ nor โ€˜bro-lyโ€™. Thank you.(endfootnote) who formulated the relationship between a particle’s wave function’s wavelength $\lambda$ and its momentum $p.$

    $$\lambda = \frac{h}{p},$$

    where $h$ is the Planck constant. Incidentally, this is the equation better known as De Broglie’s matter wave hypothesis, stating that matter, such as electrons, possess a wave-like characteristic(beginfootnote)Do note that this same equation shows that this wave-like behaviour of large bodies such as our bodies, brains, bowling balls, tennis balls, and animals is completely and utterly negligible as we will demonstrate at the end of this post.(endfootnote). This won him the Nobel Prize, no less.

    If we rewrite this to solve for $p,$ we get

    $$p = \frac{h}{\lambda}.$$

    So, clearly, a wave’s momentum is determined by its wavelength. The smaller the wavelength, the greater the momentum. What is the wavelength? It’s the length between two peaks (or two troughs). The higher the frequency, the smaller the wavelength. Now have a look at Figure 8 again. As you can see, the infinite wave of a free particle has a well-defined wavelength. The logical conclusion is that the momentum is also well-defined. Nevertheless, Figure 8 also shows that the particle’s position is not defined at all!

    Let’s turn this on its head and limit the range of possible positions of our particle. No longer is it a free particle. It is now confined within a finite range of locations.

    Figure 9. Our former free particle’s position is now restrained between $x = 0$ and $x= \pi.$ In other words, $\Delta x$ is now limited to $\pi$ wide. There is no well-defined wavelength as the wave function has different values in different places. It’s there, but not as well-defined as in the wave function in Figure 6.

    What we’ve done in Figure 9 is making $\Delta x$ smaller than it was in Figure 8 (where it was infinitely large). In fact, $\Delta x = \pi$ wide. By the same Fourier transform mechanism as with the time-frequency pair, the complimentary sister of position space $\Delta x$, namely momentum space $\Delta p_x$, will now become less certain.

    To construct a limited wave function such as the one in Figure 9, Fourier analysis shows that you need โ€“ again โ€“ a bunch of waves at different frequencies in superposition (added on top of each other).

    Figure 10. A Fourier deconstruction of the wave function in Figure 9. Many waves, many frequencies. Hence, the momentum is less well-defined.

    So, when it comes to quanta, Heisenberg’s uncertainty principle states that there’s a fundamental trade-off between information on position and momentum(beginfootnote)Another pair is energy and time. This is interesting in the context of Hawking radiation. We’ll get to that, don’t worry.(endfootnote). This is due to the fact that they are a Fourier transform pair or conjugate variables.

    This also means that if you constrain a particle to a minuscule $\Delta x,$ its wave function will start to contain momenta $\Delta p_x$ all over the place. It will occupy many more velocity possibilities, including the much faster velocities. If you were to subsequently perform a measurement, the probability of finding it moving at higher speeds is now much larger!

    Scale and effect

    At the scale of the big bad world, we never see this effect. If you would confine a bowling ball in a limited space, you will not see its momentum increase dramatically. It won’t suddenly start bouncing up and down. Conversely, if you swoop the bowling ball with considerable momentum, it won’t suddenly start appearing everywhere and nowhere at the same time: its position is still quite clear. You won’t suddenly quantum tunnel through the pins or be rolling on all bowling lanes of the neighbouring players at the same time. If it doesn’t hit a single pin, then that’s not because it’s suddenly in a state of superposition with regard to its possible locations of existence. You’re just not that good.

    You won’t notice any of these quantum effects in your everyday-scaled objects. Only when you’re dealing with particles. Or atoms. However, as soon as the mass increases, it all changes. Why? Partly because Planck’s constant is so darn small(beginfootnote)And because the number of interactions between atoms increase exponentially, causing any quantum effect to disappear due to decoherence.(endfootnote). It’s just $5.273 \times 10^{-35} \text{ J/s},$ remember? That’s small.

    All this knowledge does allow for some fun calculations. For instance, if you were to confine a bowling ball with a mass of $7.2$ kg (16 lb) inside a box where $\Delta x = 22$ cm (8.66 inches), by Heisenberg’s uncertainty principle, the ball’s speed will be $3.283 \times 10^{-35} \text{ m/s}.$ That means that after $965.9$ billion years it might have moved a distance equal to the diameter of a proton. That amount of time is seventy times the age of our current universe. Granted, quantum-mechanical effects aren’t zero, but as you can see (or rather, as one can calculate), on our everyday scale, these effects are quite meaningless.

    Sometimes, weird films such as What the #$*! Do We (K)now!? and What the Bleep!?: Down the Rabbit Hole will want to make you believe such quantum things can happen anyway. They will mention Heisenberg’s uncertainty principle like it is a magical law allowing us to do whatever. I hope that this post has shown that Heisenberg’s uncertainty principle is not about that. Nor does the uncertainty principle itself have its roots in quantum mechanics. It’s basically wave mechanics, the classical stuff, which all first-year undergraduates in physics have to learn in their first or second semester.

    A few months ago, I stumbled across a video showing an Australian senator’s question to the head of the Commonwealth Scientific and Industrial Research Organisation, an Australian federal government agency responsible for scientific research. Clearly, the senator had โ€“ shall we say โ€˜read something about Heisenberg’s uncertainty principleโ€™. During a senate hearing for a legislative committee, the senator questioned if research done in climate change should be taken with precaution as Heisenberg’s uncertainty principle stands in the way of accurate measurements(beginfootnote)He basically sought a โ€˜scientificโ€™ way to put climate science in doubt โ€“ which, apparently, he is not a proponent of. I do not claim to know anything about Australian politics, or even at great depth about climate science, however, when a legislator starts talking quantum physics โ€“ well, I do know stuff about that.(endfootnote).

    I suspect this discussion pertained to a study where a satellite uses infrared radiation to perform surface and/or atmospheric remote sensing. He continued to state that as infrared light has lower frequencies than visible light, it’s โ€˜very difficultโ€™ to understand the properties of infrared radiation based on Heisenberg’s uncertainty principle.

    Many things were going on (wrong) in this one short bit of speaking time of the senator, as is usual when someone hasn’t caught up on quantum physics as much. Which is understandable, but no less gnawing to watch (the link opens a new tab and leads to a short video on Twitter).

    In any case, I genuinely hope that this article contributed at least a sliver of knowledge to educate the electorate of the world, so we can all vote as informed and responsible as possible for the right persons for the right jobs, besides one’s preferred socioeconomic idealism.

    If you should take one thing from this post, it’s that Heisenberg’s uncertainty principle is not about anything spiritual nor does it have anything to do with scientific measurement mistakes: it’s good, old wave mechanics and Fourier analysis taught to undergrads in their first year at university. It works and it works well. It does not lead to science not being able to know things about the universe. In fact, it increased our knowledge of it. In fact, no modern information device would have worked without it. After all, you’re reading this with an electronic device which exists thanks to Fourier, Heisenberg, and De Broglie, among others. All that with a bit of more maths and more physics at the same time.

    Photo Werner Heisenberg by Friedrich Hund, a German physicist who took this photo in Heisenberg’s place of residence, Gรถttingen, in 1924. It was uploaded to Wikimedia Commons under CC BY 3.0 by Friedrich Hund’s son, Gerhard Hund, a German mathematician, computer scientist, journalist, and chess player. We have used a colour-corrected version by Martin Geisler.

  • Quantum entanglement: the EPR paradox and Bell’s Theorem

    Quantum entanglement: the EPR paradox and Bell’s Theorem


    When the state of a subatomic particle cannot be described by a wave function without taking the state of another subatomic particle into account, we speak of quantum entanglement. It’s the special case where both particles can only be described by one and the same wave function. No longer are they separate entities nor do they have separate wave functions. The astonishing consequence is that performing a measurement on one particle has an immediate effect on the measurement of the other particle, no matter how far apart they are from each other. In this article, the second part of our mini-series on quantum entanglement, we will discuss the EPR paradox which Einstein and colleagues put forward. After that, we will discuss Bell’s Theorem which allowed physicists to test Einstein’s proposal. Was Einstein correct?

    A representation of an electron’s spin โ€“ do note that this is not what an electron actually looks like nor is it what its spin looks like. The quantum world is simply too strange to depict accurately using ‘classical’ notions as done here. Here we drew a vague ball-like thing which seemingly spins around, which it isn’t and it doesn’t. But it’s the best we’ve got. Although, the best we’ve got is actually something else: a mathematical expression, the wave function.

    Quick summary

    Firstly, let me give a quick summary of the previous post:

    1. we used the property of spin as a way of distinguishing between the two entangled electrons;

    2. the orientation of an electron’s spin is expressed as spin up (anticlockwise) or spin down (clockwise) along the axis of measurement;

    3. you can arbitrarily choose along which axis you want to measure its spin, in three dimensions;

    4. no matter which axis you choose, the result is always going to be a spin up or spin down (there is no spin-a-bit-to-the-right, for instance);

    5. we are able to entangle particles in such a way that they will either always yield opposite spin or they always yield identical spin; once prepared this way, they will never deviate from this correlation when measured;

    6. we used the opposite-spin entanglement in our example and we will do so again here;

    7. quantum mechanics states that before measurement neither electrons have a specific spin: the wave function contains all possible measurement outcomes, in this case pertaining to both spin up and spin down (which can be characterised as having no definite spin yet)(beginfootnote)Analogously, the double-slit experiment showed that before measurement, particles don’t have a specific location yet.(endfootnote);

    8. as soon as you measure one electron’s spin along a certain axis, the other electron’s spin immediately snaps to the opposite orientation along that same axis, regardless of spatial distance between the two entangled particles(beginfootnote)Or, if their entanglement were prepared in such a way that they always have identical spin, the other electron would then immediately snap to the identical spin orientation along the same axis of measurement.(endfootnote).

    EPR paradox

    Even though Einstein understood quantum mechanics like few others, and while accepting these predictions and results, he didnโ€™t quite like the non-local implications brought forth by quantum entanglement. He didn’t like point 8 of the previous section. There seems to be zero time delay between influencing a particle in Amsterdam (through measuring its spin) and influencing its entangled particle in Boston. It violates a pivotal consequence of Einstein’s theory of special relativity: no signal or piece of information โ€“ anything within this universe, really โ€“ can exceed the speed light(beginfootnote)In a vacuum.(endfootnote) or else causality would not exist. In other words, if information or signals were able to travel faster than light, an effect could occur before its cause had taken place. To put it mildly, this doesn’t seem to be the universe you and I are living in.

    So, Einstein, Podolsky, and Rosen (EPR) hypothesised that something else, something secretive was going on in nature โ€“ well out of sight for theoretical and experimental physicists. Quantum mechanics as it was known then had to be incomplete. Obviously, they acknowledged its successes, but when it came to quantum entanglement, they asserted something was missing in the theory of describing nature through wave functions.

    To solve for the seemingly faster-than-light signal, they proposed that what really was going on was that the particles have always been in a specific state. When the electron pair were separated from each other, they have always had either spin up or spin down from the start from the moment of their creation.

    Suppose, a pair of gloves were made. Like all pairs of gloves, they always were each other’s opposite with respect to โ€˜handednessโ€™(beginfootnote)โ€˜Handednessโ€™ in this context is a form of the more generalised term chirality.(endfootnote). One has always been left-handed, the other has always been right-handed. And if the first one happened to be right-handed, then the other was left-handed. (Or else you’re holding a glove from another pair.)

    Suppose, the machine which had made the pair put each glove in a separate box. We can’t see which glove went in which box until we open the box. The boxes were sent to Amsterdam and Boston. The experimental physicists then open the box in Amsterdam: it’s the right-handed one! And so, we now instantly know, the one in Boston is left-handed. No magic, no non-locality, no lightspeed-breaking shenanigans.

    This is what Einstein and friends said was happening in the case of electrons. An electron pair always had specific spins to start with. It’s only in Amsterdam and Boston that we ‘open the box’ aka measure their spin. It’s only logical now that as soon as you know which spin the Amsterdam electron has, you immediately know which spin the Boston electron has.

    So, said Einstein, non-locality is an illusion. It’s all just normal local laws of nature and a bit of logical thinking. For one, spin orientation is merely hidden from us and not principally uncertain. Secondly, there’s no spooky action at a distance[1], as he famously described it(beginfootnote)In German, he wrote ‘spukhafte Fernwirkung'[1].(endfootnote).

    In everyday parlance, physicists call this a local version of the ‘hidden variables’ theory. ‘Hidden variables’ pertain to the stuff that we can’t see yet (such as spin orientation or other variables influencing this) because our quantum mechanical description (the wave function) is incomplete, however, they are there, they do exist โ€“ they do not not exist yet, according to the hidden variables theory.

    Bell’s inequalities

    Unfortunately, Albert Einstein passed away in 1955. And Niels Bohr, the other great physicist with whom he used to debate the fundamental nature of quantum mechanics passed away in 1962. In both cases too soon for them to be able to read John Stuart Bell’s 1964 paper called ‘On the Einstein Podolsky Rosen Paradox'[2]. Bell realised that Einstein’s proposal was in principle testable. It yielded a clear prediction, called Bell’s inequality.

    At this point, we must note that over the years, more than one Bell’s inequalities have been put forward by physicists(beginfootnote)Besides his original inequality, there’s the much-used CHSH-inequality, for instance.(endfootnote). To explain Bell’s inequality, we will apply a version of David Merminโ€™s original version as mentioned in his fantastic Boojums All the Way Through: Communicating Science in a Prosaic Age[3].

    Recall from point 3 before that we can measure an electron’s spin orientation along any axis. We’re going to be measuring along three axes. These axes will be at an angle of 120ยฐ relative to each other.

    The first axis will be the spin orientation along the vertical axis, which we will denote with the following symbols for spin up and spin down:

    $$\uparrow \downarrow$$

    The spin orientations up and down will also be measured along this second axis:

    $$\nwarrow \searrow$$

    And the spin orientations along the third axis will be denoted by:

    $$\nearrow \swarrow$$

    So, imagine two entangled electrons being separated in space from each other. The usual quantum-mechanical description of each electron is that they are in a superposition of spins up and spins down for all three axes.

    Except, Einstein says, no, no, not really: hidden behind the ‘veil of superposition’ they are in fact already in definite, specific spin orientations for each of the three axes. We just don’t yet know which until we measure them!

    He says, the electron in Amsterdam may already be in the specific spin states as follows:

    $$\left( \uparrow \searrow \swarrow \right)_A$$

    So, along axis 1 it’s spin up, along axis 2 it’s spin down, and along axis 3 it’s also spin down.

    Einstein continues and says that the entangled electron in Boston has to already be in the opposite states:

    $$\left( \downarrow \nwarrow \nearrow \right)_B$$

    And so, Einstein concludes, as soon as you actually perform a measurement in Amsterdam along the first axis, of course, you get the opposite spin in Boston. Only logical!

    Bell’s insight was that if you would work out this entire argument for all possible combinations, you could actually get a prediction of a ratio of outcomes. Here’s how that goes.

    First of all, if you measure along axis 1 in Amsterdam, that doesn’t mean you have to measure along that same axis in Boston. You could just choose to measure along axis 3. So, with the two examples above, your results would simply be that in Amsterdam you get spin up and in Boston you also get spin up:

    $$\left( \uparrow \right)_A \text{ and } \left( \nearrow \right)_B$$

    Bell then argued, if you would count the number of times you would get the combinations up-up, down-down, and of course up-down and down-up like this, you should get ratios of these combinations which should match experiment. If, however, these ratios don’t appear in the experiments, then Einstein’s hypothesis is incorrect. In that case, something entirely different is going on. The electrons were not already in a specific state, which in turn means that the non-local measurement effect in quantum entanglement does exist!

    Bell’s theorem

    So, let’s put them all together. Let’s first take our example above:

    $$\left( \uparrow \searrow \swarrow \right)_A \text{ and } \left( \downarrow \nwarrow \nearrow \right)_B$$

    If you measure along axis 1 in Amsterdam and along axis 1 in Boston you get spin up, spin down. If you measure along axis 1 in Amsterdam and along 2 in Boston, you get spin up, spin up. And so on, and so forth! We’ve put it in a little table:

    Here you can see all the possible combinations of measurement outcomes along the three possible axes of the electrons in Amsterdam (A) and Boston (B). We used U for spin up and D for spin down.

    Bell then says that if Einstein was correct, and the states of the spin orientations along these three axes were already there, then these are the expected outcomes.

    Let’s focus on the number of UD or DU combinations, in other words, let’s focus on the number of times we find the opposite spin orientations, irrespective of the axes along which they are measured. We’ve marked them yellow.

    Exactly five out nine times you will find the opposite spin directions.

    Let’s check for other spin combinations. Suppose, the electron in Amsterdam is secretly in the following spin states, $\left( \downarrow \nwarrow \swarrow \right)_A$, and the electron in Boston is then the opposite, $\left( \uparrow \searrow \nearrow \right)_B$. If we count again the number of times the measurement outcome of opposite spins, we get, again, five out of nine.

    Okay, I think you can imagine where this is going. We’re not going to go by all the tables, but I do want to do one more, just for fun. Suppose, the one in Amsterdam is all spin down, $\left( \downarrow \searrow \swarrow \right)_A$, and, obviously, the Boston one is its opposite, $\left( \uparrow \nwarrow \nearrow \right)_B$. In that case, we would get opposite spins in nine out of nine times.

    And so, this particular Bell inequality states that the probability (P) of finding opposite spins along all three axes is at least $\frac{5}{9}$ or 55% (and at most 1 or 100%). In other words, $P(\text{opposite}) \geq \frac{5}{9}$. If this inequality were violated by experiment, the underlying theory will have been proven to be incorrect.

    Experimental outcomes

    Over the past thirty years, many experiments were carried out to test multiple versions of Bell’s inequality. Usually, these tests involved photons rather than electrons and pertained to measurement of polarisation rather than spin.

    Freedman and Clauser did the first Bell test. They used a version of the so-called CH74 inequality[4].

    The most well-known test was performed by Alain Aspect and colleagues. As Bell had originally suggested, they were able to have the two measurement devices randomly select the method of measurement before the entangled photons had arrived[5].

    In all tests, all versions of Bell’s inequalities were violated. Instead, the statistical outcome was congruent with the predictions of quantum mechanics. The conclusion has to be that Einstein’s local hidden variable theory was incorrect. There is nothing local about measuring entangled particles.

    In our particular inequality, the result was that the occurrence of opposite spins turned out to be exactly 50%, not 55%.

    Conclusions

    Let’s summarise what we have established over the course of the last two posts, including this one.

    In quantum mechanics, particles which have not been measured yet don’t have a definite, specific state. Instead, they are best described by a wave function which incorporates all the possible future states it can snap into once measured.

    When a particle can only be described in tandem with another particle, i.e. both particles can only be described by one and the same wave function, they are maximally quantum entangled(beginfootnote)In practice, in the real world, particles aren’t maximally entangled like the way we can prepare them in the laboratory. The world is too messy for those ‘pure states of entanglement’ to exist for any significant amount of time. There are simply too many particles around to not interact with any other particle. Every particle will invariable interact with thousands of trillions of other particles and so any previous entanglement will quickly decohere into either a very weak version of the original entanglement or simply to zero entanglement. Every interaction represents a measurement. Since our brains are too large and consist of thousands of trillions of particles, they will never be in a pure state of superposition nor entanglement. Not to mention our much larger body, which will never be in any sort of quantum state. It is statistically so unlikely that you’d have to become as old as $(10^{100})^{100}$ times the age of our current universe to witness such an event. And that number was a metaphorical one. It’s much larger.(endfootnote).

    If their entanglement entails their spins will always correlate in a certain way โ€“ be it identical spins or opposite spins โ€“ a measurement on one particle, causing it to snap into one of the possible, specific, definite states, has immediate effect on the state of the other particle: it instantly snaps out of its wave function haze into a correlating, specific, definite state.

    Einstein didn’t like this as this would imply some kind of information was somehow transported beyond the speed of light from one particle to the other.

    He postulated that particles have always been in a specific, definite state to begin with. The only reason we don’t know which is because we haven’t measured it yet. There is no ‘snapping out of the haze’ going on.

    John Bell showed that Einstein’s hypothesis can be tested. If you would perform many, many measurements of many, many maximally entangled particles, eventually, the occurrences of the variety of correlated states should show up in a certain ratio, an inequality, as it happens.

    Experiments showed they do not. Instead, the ratio is exactly according to the predictions of quantum mechanics.

    This demonstrated that particles indeed snap out of their haze upon measurement and not that particles had always been in a hidden but definite state.

    And if that is true, then non-locality has to be true โ€“ there is no other way the other particle snaps into the correct, correlated state.

    Nobody knows how this happens. Certain non-local but still hidden-variables hypotheses have been proposed. One of the more famous versions is called the ER=EPR conjecture by Juan Maldacena and Leonard Susskind. Perhaps we’ll dive into that later on.

    Einstein’s aversion to this ‘particles have no definite state until measured upon’ made him utter his famous complaint, ‘God does not play dice’.

    Unfortunately, he was wrong here on two occasions. God(beginfootnote)We are using the word ‘God’ in a purely metaphorical way. This does not pertain to any specific religious entity as revered by many in a variety of societies in human culture.(endfootnote) does play dice. Moreover, He throws them where we can’t see them. Even God seems to be bound by Heisenberg’s Uncertainty Principle. But that’s a subject for another bit of maths and physics.


    [1] Einstein, A., Podolsky, B. and Rosen, N. (1935) โ€œCan Quantum-Mechanical Description of Physical Reality Be Considered Complete?,โ€ Physical Review, 47(10), pp. 777โ€“780. doi: 10.1103/PhysRev.47.777.

    [2] Bell, J. S. (1964) โ€œOn the Einstein Podolsky Rosen Paradox,โ€ Physics Physique Fizika, 1(3), pp. 195โ€“200. doi: 10.1103/PhysicsPhysiqueFizika.1.195.

    [3] Mermin, N. D. (1990) Boojums all the way through : communicating science in a prosaic age. Cambridge England: Cambridge University Press.

    [4] Fry, E. S. and Thompson, R. C. (1976) โ€œExperimental Test of Local Hidden-Variable Theories,โ€ Physical Review Letters, 37(8), pp. 465โ€“468. doi: 10.1103/PhysRevLett.37.465.

    [5] Aspect, A., Dalibard, J. and Roger Gรฉrard (1982) โ€œExperimental Test of Bell’s Inequalities Using Time-Varying Analyzers,โ€ Physical Review Letters, 49(25), pp. 1804โ€“1807. doi: 10.1103/PhysRevLett.49.1804.

    Featured image: Portrait of theoretical physicist John Bell at CERN, June 1982 (CERN, CC BY 4.0)

  • Quantum entanglement: non-locality and the state of a two-particle system

    Quantum entanglement: non-locality and the state of a two-particle system


    To this day, quantum entanglement and its effects are phenomena which still leave physicists scratching their heads when trying to get a deeper understanding of what is actually happening. This series on quantum entanglement is going to be a two-parter. In this post, we will discuss what is meant by locality and non-locality and what quantum entanglement is. The term quantum entanglement has been used in many instances of popular culture pertaining to spirituality, healing, and a flurry of new age approaches to human consciousness. This is not the kind of โ€˜quantum entanglementโ€™ we will discuss here. We will purely look at the physics of it, its original and proper meaning. We will study the state of a two-particle system. In the next post, we will discuss what Einstein and his friends proposed, what Bell wrote, and whether Einstein was right. And then there are also exciting caveats which we will explore.

    The basics

    Letโ€™s go over the basics one more time. โ€˜Particlesโ€™ arenโ€™t particles in the classical sense at all โ€“ theyโ€™re absolutely not like tiny balls or pellets. They are best described by the wave function, a mathematical expression containing all possible states the particle can be in. This pertains to its energy levels, its positions or a number of other properties it can have.

    As long as no measurements have been performed on it, the particle has no definite state or states. It displays wave-like behaviour like being caught in a haze of all possible states. However, as soon as you measure it, the particle will snap out of its haze and it will appear to be a particle, an actual particle in the classical sense, with a definite state.

    Note that โ€˜the state of an electronโ€™ can refer to a particle with no definite set of states when no measurement was performed. The state of an electron is then best described by the wave function, which contains all possible definite states upon measurement.

    Hereafter, โ€˜wave functionโ€™ and โ€˜stateโ€™ are used interchangeably.

    In This is not an atom, the wave function is discussed. In The double-slit experiment, the wave-like and the particle-like behaviours are showcased.

    Locality vs non-locality

    Isaac Newton knew he had a problem when he formulated his theory of gravity. While it beautifully described the extent to which two masses exert gravitational forces upon each other, his theory didnโ€™t explain how they did that. He didnโ€™t like the conclusion that the gravitational influence between Earth and the Moon seemed to spookily operate at a distance through the vacuum. He wrote it was โ€˜so great an Absurdity that I believe no Man who has in philosophical Matters a competent Faculty of thinking can ever fall into itโ€™. He famously stated to leave this unsolved mystery to โ€˜the Consideration of my readersโ€™[1].

    In other words, Newton wasnโ€™t big on non-locality. And yet, his own theory did entail an invisible force operating over vast distances through the vacuum. Moreover, it seemed to be an instantaneous effect: if the Sun were to suddenly disappear, then Earth would be flung off its trajectory immediately. Of course, today, we know that nothing can travel faster than light, so the gravitational changes of the Sun would take about eight minutes to โ€˜reachโ€™ Earth.

    The following years, physical phenomena such as magnetism and electricity proved, in fact, to be very local indeed. It became clear there is always an indirect way through which one object is able to influence another object at a distance. What is meant with locality? Hereโ€™s the mechanism: an object interacts with its immediate environment, a field embedded within the three-dimensional space we live in, i.e. the electromagnetic field, which then passes on that ripple of disturbance onto the other object. In terms of โ€˜fieldsโ€™, one could say that at one particular location the fieldโ€™s value is changed by some object. That value change then changes the values of the field in the direct vicinity, which then change the values in their vicinity, and so on. Itโ€™s a bit like โ€˜the waveโ€™ done by thousands of sports fans in a stadium. Or like falling dominoes. Every change is ever local and the propagation of that change through space is limited to the speed of light.

    Tumbling telephone boxes are definitely a โ€˜local phenomenonโ€™. The sculpture Out of Order by David Mach is situated in Kingston upon Thames (UK). Photo by 272447.

    Many years later, Einstein replaced Newtonโ€™s theory with his own theory of gravity, General Relativity (GR). It showed that Newtonโ€™s intuition was correct. Gravity couldnโ€™t be non-local and Einstein showed it isnโ€™t. In GR, space and time itself are the stretchy substance through which gravitational disturbances propagate at the speed of light towards the other object. When a mass curves or disturbs spacetime around it, that curvature or disturbance then ripples through the universe, on its way to influence other objects. In fact, on 11 February 2016, a large collaboration of incredibly talented scientists physically measured these gravitational ripples in spacetime as predicted by Einstein in 1916. It won three key figures the Nobel Prize.

    And so, it seems there is no spooky influence at a distance in physics. Even still to this day, in modern quantum physics, our best understanding and most successful theory is that quantum fields pervade our universe, forming the mediums through which forces are propagated, limited by the speed of light.

    Non-locality entails a change in one patch of space instantaneously influencing another patch of space irrespective of their distance. Locality entails the propagation of change through space by influencing only neighbouring patches of space at a maximum of the speed of light.

    Spin

    Electrons have several properties. One of the more obvious is (negative) charge. The Stern-Gerlach experiments showed that they possess another property which was given the name spin angular momentum or simply spin for short, for lack of a better term as electrons arenโ€™t exactly like spinning balls.

    Nevertheless, as it stands, electrons have an intrinsic spin, which cannot in any sensible way be described like a classical-mechanical rotation. Like with any object in three-dimensional space, you can measure its spin along any angle within 360 degrees in three dimensions. With respect to whichever axis you choose, they can only ever spin clockwise or anticlockwise(beginfootnote)Yes, this does sound like there is an actual rotation around an axis in the classical sense. And maybe, in some deep sense, there is after all, however, this deserves a post of its own, so suffice to say for now, our language is simply too limited to avoid using classical terms for quantum mechanical phenomena, misleadingly.(endfootnote). The latter is called spin up and the former spin down, according to the right-hand rule.

    If electrons were like tiny, fluffy balls such as displayed here, you could picture their spin as an anticlockwise or clockwise rotation about the axis of measurement. Using the right-hand rule, we can designate this spin-up or spin-down. Of course, in three dimensions, any axis of measurement at any angle can be chosen with respect to which it will be found spinning. Disclaimer: this classical-mechanical illustration does not portray actual electrons nor actual quantum mechanical spins. But itโ€™s perhaps useful as a simile. (Illustration by KJ Runia)

    Symbols

    As we take our readers seriously, weโ€™ll take this opportunity to introduce a few mathematical symbols which will prove to come in handy at later stages of this series.

    Letโ€™s use the symbol $\lvert A \rangle$ to denote the state of the electron in Amsterdam with respect to its spin. As long as we havenโ€™t performed any measurements on the electron, it has no definite state. However, upon measurement, its spin with respect to the vertical axis of measurement is ever either spin up or down. Letโ€™s write these two possible measurement outcomes as $\lvert\uparrow\rangle_A$ or $\lvert\downarrow\rangle_A$.

    Likewise, if the state of an electron in Boston $\lvert B \rangle$ is spin up or spin down, we write $\lvert\uparrow\rangle_B$ or $\lvert\downarrow\rangle_B$.

    Assuming the state of the electron in Amsterdam hasnโ€™t been measured yet, we can express this (with respect to spin) as a combination of both spin states:

    $$\lvert A \rangle = \alpha \lvert\uparrow\rangle_A + \beta \lvert\downarrow\rangle_A .$$

    This is why physicists often poetically say that the unmeasured particle is in a state of both spins at the same time while itโ€™s more accurate to say it has no definite state. Mathematically, its state is an amalgam of all possible, linearly superposed (added together), algebraic solutions to the Schrรถdinger equation, hence, itโ€™s said to be in quantum superposition.

    Whatโ€™s that $\alpha$ and $\beta$, you ask? Well, theyโ€™re numbers of probability we need to find in order to complete our expression. The Born rule states that if we square the (modulus of the) wave function (the state), we will get the probability (density) of either possible outcome after measurement. Now, experiments have shown that either outcome, spin up or spin down, $\lvert\uparrow\rangle_A$ or $\lvert\downarrow\rangle_A$, appears in 50% of the total number of measurements. In other words, the probability of measuring either spin state is exactly $\frac{1}{2}$. So, if we put $\alpha=\beta=\frac{1}{\sqrt{2}}$, then $\lvert\alpha\rvert^2 = \lvert\beta\rvert^2 = \frac{1}{2}$. After all, $(\frac{1}{\sqrt{2}})^2 = \frac{1}{2}$, which is exactly what we want. So, the state (wave function) of our Amsterdam electron with respect to spin can be represented by

    $$\lvert A \rangle = \frac{1}{\sqrt{2}} \lvert \uparrow\rangle_A +\frac{1}{\sqrt{2}} \lvert \downarrow\rangle_A .$$

    Similarly, the state of the electron in Boston with respect to spin is then represented by

    $$\lvert B \rangle = \frac{1}{\sqrt{2}} \lvert \uparrow\rangle_B +\frac{1}{\sqrt{2}} \lvert \downarrow\rangle_B .$$

    What you need to take from this is the following: the state of an electron before measurement is the sum of all possible states (multiplied by a probability factor, in this case $\frac{1}{\sqrt{2}}$).

    In the case of spin as measured along the vertical axis, the state of the electron is the sum of two possible states, spin up $\lvert \uparrow \rangle$ or spin down $\lvert \downarrow \rangle$.

    Note that there are other possibilities: we could measure the spin along a horizontal axis. We could represent this with spin left $\lvert \leftarrow \rangle$ or spin right $\lvert \rightarrow \rangle$. Or we could measure the spin at angles of +120 or -120 degrees from the vertical axis, which we might represent as $\lvert \nwarrow \rangle$ and $\lvert \searrow \rangle$ or $\lvert \nearrow \rangle$ and $\lvert \swarrow \rangle$. We will get to that in the discussion of Bellโ€™s Theorem in the next post.

    Quantum entanglement

    So, what is quantum entanglement? Recall that the most complete description of a particle is the wave function. This has always been about a free, single particle, not interacting with anything. In the case of quantum entanglement, however, this doesnโ€™t fly anymore.

    When the state of a particle can no longer be described without a description of the state of another particle, those two particles are said to be quantum entangled. No longer can we describe either particle by one wave function each. They can only be described as a two-particle system by one and the same wave function.

    This has an astonishing consequence. Suppose our two electrons become entangled in such a way that they always have opposite spins(beginfootnote)Producing spin-entangled electrons is difficult but clever experimental physicists have their ways.(endfootnote). So, if one has โ€˜spin upโ€™, $\lvert \uparrow \rangle$, the other always has โ€˜spin downโ€™, $\lvert \downarrow \rangle$, or vice versa(beginfootnote)Itโ€™s also possible to have them correlate such that they have identical spin, but for our example, letโ€™s not.(endfootnote). So, we now have one system with two particles who always have opposite spins, which means that the total spin of our system is 0, zero. Letโ€™s denote the total spin of our system with $\lvert S \rangle$.

    Before our experiment takes place, they are both separated. One is staying in a laboratory in Amsterdam. The other is transported to Boston. Since no measurement has taken place on either particle, they are in a superposition according to the one wave function. They havenโ€™t an exact location (although one is very likely to be somewhere in Amsterdam at the moment of measurement and, likewise, the other in Boston), their energy levels are all over the place, and their spin isnโ€™t either spin up or spin down along this or that axis.

    We can represent this whole situation with respect to spins as follows:

    $$\lvert S \rangle = \dfrac{1}{\sqrt{2}} \left( \lvert \uparrow \rangle_A \lvert \downarrow \rangle_B – \lvert \downarrow \rangle_A \lvert \uparrow \rangle_B \right) .$$

    When youโ€™re looking carefully at the expression above, you can see that the state of the total spin $\lvert S \rangle$ of our two-particle system is a combination of two situations: the electron in Amsterdam is spin up and so the electron in Boston is spin down or the electron in Amsterdam is spin down and the electron in Boston is spin up. They need to be subtracted from each other because the total spin equals 0, remember? Hence, the minus sign. Lastly, both states are multiplied by the fraction $\frac{1}{\sqrt{2}}$ because both states have a 50% chance of occurring (which you get if you square the whole thing).

    And so, what does this mean? As soon as you perform measurements on the one in Amsterdam, and you find it has spin up, the other electron in Boston immediately has spin down along that particular axis upon measurement, even though the probability before measurement was still 50%! How does the electron in Boston โ€˜knowโ€™ what the measurement result in Amsterdam was? En how does it know this so fast? Faster than the speed of light! Besides this, turns out, youโ€™ll always get a definite spin from the other particle opposite to the one you measured first. As soon as the measurement in Amsterdam took place, the measurement outcome in Boston being the opposite result is always 100% all of a sudden! (Or the other way around.) There are never any exceptions!

    In other words, as soon as you do the measurement, the mathematical description changes from

    $$\lvert S \rangle = \dfrac{1}{\sqrt{2}} \left( \lvert \uparrow \rangle_A \lvert \downarrow \rangle_B – \lvert \downarrow \rangle_A \lvert \uparrow \rangle_B \right) ,$$

    to either

    $$\lvert S \rangle = \lvert \uparrow \rangle_A \lvert \downarrow \rangle_B ,$$

    meaning, the state of the total spin equals the one in Amsterdam being spin up and the one in Boston being spin down, or, vice versa:

    $$\lvert S \rangle = \lvert \downarrow \rangle_A \lvert \uparrow \rangle_B .$$

    And hereโ€™s the astonishing part: this will always work this way, no matter how great the physical distance between the two particles. Locality out the window. Welcome back, non-locality.

    Einstein accepted this prediction in quantum mechanics as being correct. However, he didnโ€™t like it. How did the other particle instantly โ€˜knowโ€™ which spin to exhibit when Einsteinโ€™s fantastically successful theories of relativity relied on the universal law that nothing can exceed the speed of light? He accepted the theory but he concluded it wasnโ€™t complete. There had to be some sort of hidden mechanism which they had overlooked.

    We will discuss Einsteinโ€™s attempt at saving the principle of locality and the universal speed limit in the next post. As well as John Bellโ€™s and Alain Aspectโ€™s subsequent work. For now, the question of whether Einstein was right, we will โ€˜leave up to the Consideration of our readers.โ€™


    [1] Newton, I. (1756) Four Letters from Sir Isaac Newton to Doctor Bentley: Containing Some Arguments in Proof of a Deity [Online]. Available here. (Accessed: 14 May 2020)

    Featured image by KJ Runia

  • Lab centrifuges and prime numbers

    Lab centrifuges and prime numbers


    When micro- or molecular biologists do research on viruses, bacteria, fungi, human or animal cells, one of the many instruments they will use is a laboratory centrifuge. This equipment allows them to separate substances contained within a test tube. This way scientists are able to obtain, for instance, purified enveloped viruses, such as the novel coronavirus, SARS-CoV-2. Or they can isolate nucleic acids, such as DNA.

    Often, the rotor of the machine rotates at incredible speeds. It is vital that the test tubes have been placed in a perfectly balanced way. If not, the machine might break down and potentially dangerous glass shards and substances might be flinging about(beginfootnote)Although sensors may be installed to prevent the machine from operating in case of force imbalance. See also the Final remarks down below.(endfootnote).

    Fortunately, there is a nifty way to calculate whether you can โ€“ in principle โ€“ place a certain number of test tubes in an evenly balanced way. To crack the code, we will use my favourite type of number: the prime numbers. Fun fact: this funky little trick wasnโ€™t proven until fairly recently in 2010.


    NEW: Listen to the audio |


    The set-up

    Before we begin, we assume that the mass of each test tube, including their contents, is equal. Also, I would like to remark that, of course, we could do this the physics way, using angular velocity and torque and all that, but in this case, weโ€™re going to be all mathy about it, or specifically, in a way, number-theoretical.

    Suppose, the machine can hold eight test tubes. Eight holes are positioned in a circle on the rotor bit of the machine.

    If we have just one test tube, thereโ€™s no way we can make it balanced. That much is clear. If we have two test tubes, however, no problem. They can be balanced easily. Just put one on either side precisely opposite each other. Three test tubes? Hm. I donโ€™t see how. Whatever arrangement we try, itโ€™s always going to be asymmetrical. What if you have four test tubes? Well, this is easy enough. Make it symmetric, like a square.

    Okay, so what about five test tubes? Well, thatโ€™s just the same as when we had the inverse of this, with three test tubes! That couldnโ€™t be done, so, this canโ€™t be done either.

    Six? Yeah, of course, we can do that. Itโ€™s just the same as having two test tubes, itโ€™s just the inverse! Three on one side and three on the other side. Now you have two open spots on either side. Perfectly symmetrical, just like the inverse situation, where you had two test tubes and six open spots.

    Seven? No. You will have guessed it by now. Having seven test tubes is exactly the same as having just one test tube in a rotor with eight spots.

    And eight, well, of course, we can do eight. Itโ€™s also the exact same as having no test tubes at all. So, yes, thatโ€™s balanced.

    Do you see a pattern here? You might. Notice how the number of occupied spots and empty spots always complement each other.

    Prime factorization

    Just for clarityโ€™s sake, Iโ€™m going to call whole numbers integers since thatโ€™s what theyโ€™re called in mathematics.

    So, Iโ€™m assuming we all know what a prime number is: an integer greater than 1 which cannot be formed by multiplying two smaller integers. In high school or even in primary school, you may have been taught that prime numbers are numbers which can only be divided by 1 or by itself (not including 1). So, prime numbers are 2, 3, 5, 7, 11, 13, 17 and so on.

    Prime factorization is writing down any non-prime integer as a multiplication of two or more prime numbers. The fundamental theorem of arithmetic states that any integer is either itself a prime number or can be written as a product of prime numbers. This is one of the reasons why theyโ€™re my favourite. Primes are the building blocks of any integer.

    So, for instance, we take the number 15. This number can be written as $ 15 = 3 \times 5 $. Or take 279. We can write $ 279 = 3 \times 3 \times 31 = 3^2 \times 31 $. Letโ€™s take 16. This number can be written down as $ 16 = 2 \times 2 \times 2 \times 2= 2^4 $.

    As you can see, prime factorization is pulling apart a non-prime number into a product of prime numbers. We call the latter prime factors. 

    So, thatโ€™s what that is. One of the many applications of prime factorization is finding the greatest common divisor between two integers, for example. Or encrypting (and decrypting) secret files and messages. Here, weโ€™re going to use it for calculating whether test tubes can be arranged in a balanced way.

    The trick

    Suppose, your machine has $n$ spots available. Suppose, $k$ is the number of test tubes. The number of empty spots is $n-k$. Hereโ€™s the trick.

    Determine the prime factors of $n$. If (and only if) $k$ can be written as a sum of these prime factors and the number of empty spots $n-k$ can be written as a sum of these prime factors, you can in principle balance the rotor.

    The mathematics

    Itโ€™s too technical to discuss at length the proof given by Gary Sivek in his 2010 paper (or here). However, the gist for the more mathematically inclined is available by clicking โ€˜expandโ€™. You may skip this paragraph if this is (understandably) still too technical.

    Expand

    Striving to obtain an $n$-th cyclotomic polynomial (or prime polynomial), we obtain a series of complex numbers $z^n$ which satisfy $z^n = 1$, all being $n$-th roots of unity where $n$ is the number of total spots on the centrifuge. We then map the test tubes onto the roots of unity in a non-overlapping way. As is well-known, the values of $z \in \mathbb{C}$ are given by $e^{\frac{2\pi i}{n} k}$, where $1 \leqslant k \leqslant n$.

    So, now we have $k$ roots of unity among the $n$-th roots of unity representing the occupied spots in the centrifuge.

    Sivek proved, using Leungโ€™s and Lamโ€™s Theorem, that if (and only if) the sum of the $n$-powered $k$ roots of unity and the sum of the $n$-powered $n-k$ roots โ€˜vanishโ€™, i.e. are equal to zero (using good-old de Moivreโ€™s formula, if you remember from your very first semester at uni), as long as $n \geqslant 2$ and $1 \leqslant k l\eqslant n-1 $, then balancing is a fact (where $k=0$ and $k=n$ were regarded to be trivial cases for obvious reasons).

    As you can see, no classical mechanics required.

    An example with eight roots of unity in the complex plane

    Obvious examples

    Suppose, we take our centrifuge which was capable of handling 8 test tubes. We have 6 test tubes. First thing we do is calculate which prime factors the number 8 has. We know this, itโ€™s all 2s. So, the only prime factor of 8 is 2. We can write the number of test tubes, 6, as a sum of this prime factor 2: $6 = 2 + 2 + 2$. The number of empty spots, thatโ€™s $8-6 = 2$, is the prime factor itself! So, yes, if you have 6 test tubes, you can balance the machine.

    Letโ€™s take 7 test tubes. Can this be written as a sum of the prime factors of 8? No, it canโ€™t. Well, thatโ€™s it then. We cannot arrange the test tubes in such a way that itโ€™ll be balanced out.

    A counter-intuitive example

    Suppose, our centrifuge is capable of handling 12 test tubes in total. We only have 7 test tubes. Hm. Surely, we can imagine 6 test tubes working, but can we make a balanced arrangement with 7 test tubes?

    Letโ€™s first do some prime factorization with 12. So, $ 12 = 2 times 2 times 3 = 2^2 times 3 $. In other words, the prime factors of 12 are 2 and 3.

    Now, can we write 7 as a sum of these prime factors? Yes, we can: $7 = 2 + 2 + 3$. Okay, so far, so good. Can we write the number of empty spots as a sum of these prime factors? Well, $12-7 = 5$. And yes, we can also write 5 as a sum of 2s and 3s: $5 = 2 + 3$.

    So, yes, we can balance 7 test tubes in a rotor with 12 spots! Itโ€™s likely this outcome wasnโ€™t immediately apparent to you. If you were to see or draw a depiction and a working out of the arrangement yourself, however, I think itโ€™ll become clear how this would work. Bonus points if you can draw a balanced configuration for 5 test tubes. Because you should know by now, you can.

    Bonus trick

    The beauty of it all is that all of the above does give us another quick way to assess whether we can balance the centrifuge. Iโ€™m going to be honest with you: it may be the easiest. If you can express the number of test tubes as the sum of two numbers of which you already know you can balance the rotor, then you can balance the rotor. Heh.

    Final remarks

    In real life, most machines have sensors to prevent force imbalances from taking over. The rotors have markings so that users wonโ€™t have to think about where to place the test tubes. Besides, in a university lab, you would simply make sure you prepare the number of samples which make a balancing act trivial. Moreover, many rotors contain three compartments containing sets of test tubes. This makes adjusting for mass variability much easier. And some machines, in hospital labs, for instance, have fully automated robots doing the heavy lifting.

    Therefore, the reason for why this type of mathematics is done, isnโ€™t so much for the applicability as it is for the joy of exploring deep connections such as between prime numbers and complex geometry, if you will. Itโ€™s first and foremost a fun and fruitful exercise of human exploration of the lands of number theory, algebraic geometry, and finite fields, on the continent that is pure mathematics.

    Featured image by Michail Tzortzatos under CC BY-SA 4.0
    Spinning rotor by user musicalwoods under CC BY-SA 2.0

  • The double-slit experiment

    The double-slit experiment


    Over three hundred years ago, grumpy old men with 17th-century wigs or 18th-century black-ribboned man ponytails were divided into two camps. They were squabbling over what type of phenomenon light is. โ€˜Light is wavesโ€™, said Huygens, Hooke, Euler and friends. โ€˜No no, light is particlesโ€™, said Newton, Laplace and colleagues. Fast forward to 1990 and even my physics teacher in high school confesses he still wasnโ€™t sure about the correct answer.

    His confusion is understandable. Even though he could have known the correct way of thinking about it, the reason for these murky waters can be traced back to the now famous set of double-slit experiments.

    So, join me in tumbling through the slits of science, into the mad world of quantum physics, where one thing was proven be in two places at once. Or was it?


    NEW: Listen to the audio |


    (not really meant as podcast since referrals are made to figures in the article)


    The set-up

    Suppose you had a shotgun capable of spraying a cloud of numerous tiny lead pellets in one shot. If youโ€™d aim it at a screen containing two thin slits, so only some might get through, what shooting pattern should you expect to appear on a screen behind it?

    Figure 1. A screen with two slits
    Figure 1. A screen with two slits

    Iโ€™m quite confident your answer will correlate strongly with the situation as depicted in Figure 2.

    Figure 2. The screen with the double slits and a screen behind it with the typical impact pattern of pellets or particles
    Figure 2. The screen with the double slits and a screen behind it with the typical impact pattern of pellets or particles

    This is exactly what you would expect if the things youโ€™re using to shoot with are tiny pellets or tiny particles. No surprise here.

    Now imagine, weโ€™d slowly submerge the screen with the two slits half-way into a pond. Water waves are slowly rolling towards the first screen as depicted in Figure 3.

    Figure 3. Both screens are now partly submerged in water. Water waves are approaching the first screen.
    Figure 3. Both screens are now partly submerged in water. Water waves are approaching the first screen.

    What would these waves look like after theyโ€™ve gone through the slits? When seen from above, it would look like Figure 4.

    Figure 4. As the waves go through the two slits, they transform into two circularly spreading waves like two stones in a pond.
    Figure 4. As the waves go through the two slits, they transform into two circularly spreading waves like two stones in a pond.

    The two slits transform the waves into two circularly spreading waves. Like two stones thrown into a pond. You can see how the waves will intersect with each other. You might expect some interaction to occur at these crossroads and you would be right.

    In fact, letโ€™s have a look at a real pond. In the GIF of Figure 5, you can clearly see how these two circular waves interfere with each other. If two crests meet, they amplify each otherโ€™s amplitude, whereas two troughs meeting, they amplify each otherโ€™s trough-ness (also amplitude but in the other direction). And where a crest meets a trough, they cancel each other out!

    Figure 5. Two circular waves in an actual pond.
    Figure 5. Two circular waves in an actual pond.

    Now have a look at the animation of Figure 6 and observe especially what the second screen receives: patches where the waves hit the screen are white and patches where there are no waves at all are black.

    Figure 6. The white areas are where the (amplified) waves hit the second screen, the black areas are the parts where no waves are present due to mutual cancellation
    Figure 6. The white areas are where the (amplified) waves hit the second screen, the black areas are the parts where no waves are present due to mutual cancellation

    So, now we know what happens if waves would be thrown at the two slits. Contrary to what you see when you would shoot pellets towards the screen, you would see whatโ€™s depicted in Figure 7.

    Figure 7. The double-slit set-up with the typical pattern on the second screen when waves have gone through

    So, now we have two options. If whatever weโ€™re shooting at the slits is particles, we get whatโ€™s on the left in Figure 8. If weโ€™re aiming waves at the slits, we get what is on the right in Figure 8.

    Figure 8. If particles went through the slits, you get to see the pattern on the left. If itโ€™s waves, you get the pattern on the right.

    Youngโ€™s interference experiment

    Thomas Young was a polymath and physician. In the 1790s, he wrote a thesis on the physical and mathematical properties of sound. In 1800, he presented the Royal Society, the UKโ€™s national academy of sciences, his theory that light is a wave too. He was met with great skepticism as the likes of Newton and Laplace were proponents of the light-is-particles theory. 

    Young then showed how they were wrong. A notable fact is that he didnโ€™t actually use two slits. He had a bundle of sunlight pass through a pinhole so as to obtain a very tiny bundle of sunlight. He then placed a โ€˜slip of cardโ€™ in front of the pinhole, essentially splitting the small bundle in two even smaller bundles which then interfere with each other. The resulting light pattern would have looked like the one shown in Figure 9.

    Figure 9. The pattern which Young produced by splitting sunlight

    If light were particles, you would have seen an entirely different pattern. This result, however, completely corresponds to the wave theory of light. Young concluded therefore that light is indeed a wave phenomenon. He called this the most important of his achievements.

    This marked the beginning of the acceptance of the wave theory of light (yay for Huygens and friends) and a departure from the particle theory of light (nay for Newton and fr… well, colleagues, at least).

    Or particles after all?

    Figure 10. Individual electrons

    Of course, Max Planck, Albert Einstein, and a few other colleagues would later show that light is particles after all. In a previous post, The formula that got Albert Einstein the Nobel Prize and should stop us getting sunburn all the time, we discussed Einsteinโ€™s finding which won him the Nobel Prize.

    In short, Max Planck and Albert Einstein showed that certain behaviour of light could only be explained if it consisted of small packets of energy, quanta as they were labelled.

    But apart from that, experimenters found another peculiarity. In the 1960s, electrons were generally expected to behave like particles โ€“ like pellets or ball bearings. So, instead of light, they fired one electron at a time towards a splitter and have a screen behind that capture the electron. What they initially saw was to be expected. A few (11) loose dots on the screen as shown in Figure 10. However, as the individual electrons kept being fired, one after the other, an astonishing pattern started to emerge โ€“ the kind you would expect to see in the case of interfering waves! Wait, what?

    Are they waves after all? But they were individual electrons! How?

    Needless to say, experimenters did the same thing with individual photons, the quanta of light Max Planck and Einstein were talking about. Extremely low-intensity light was produced up to the point where single photons were shot at the screen. The same result. They seem to behave like particles at first but then this wave pattern emerges.

    Two places at once?

    Theorists then theorised that the only explanation was that a single electron and a single photon somehow went through the two slits at the same time, enabling some kind of self-interference so that this wave pattern would emerge while also preserving a particle pattern at the same time.

    To test this theory, people put particle detectors at the two slits in order to see if the single electron or the single photon indeed flew through both slits at the same time.

    The result was again astonishing: the wave pattern disappeared and what they got was instead the pattern youโ€™d expect to see if the particles were actual particles โ€“ the pattern was like the pattern in Figure 2. At the same time, they never detected the particle at both detectors. They were ever only detected by one detector โ€“ as if they were particles.

    As soon as they removed the detectors, however, the wave pattern emerged again.

    And even if they placed just one detector at one slit, the wave pattern disappeared again and the particle pattern showed up.

    It was as if the electron and the photon knew when they were being watched and then decided to behave differently.

    This is called the measurement problem. In the next post, we will discuss this at greater depth.

    People now started talking about the wave-particle duality of elementary particles. Are particles truly particles or waves? Theyโ€™re both, people now said. Sometimes theyโ€™re waves, sometimes theyโ€™re particles.

    Fields

    Of course, nowadays, the reigning theoretical paradigm is quantum field theory โ€“ mathematical field descriptions to capture the behaviour of โ€˜particlesโ€™ such as the electron, the photon, and a whole zoo of elementary constituents of our reality. The most successful quantum field theory to date is called the Standard Model of particle physics. In a previous post, Why, exactly, do glass and liquids refract light?, we dive a little bit into quantum field theory.

    In short, the question of whether light is particles or waves has been answered: itโ€™s fields. The same goes for electrons. And all the other elementary โ€˜particlesโ€™. Itโ€™s all fields.

    As long as no interaction with the outside world such as detectors take place, a photon or electron are part of the wave functions of their respective electromagnetic and electron fields, governed by the Schrรถdinger equation. They are very much like waves. However, as soon as they interact with something, such as a detector, what is detected is a particle, merely a slice of a photonโ€™s or electronโ€™s entire wave function.

    I promise we will unpack these two last paragraphs in a later post. We expounded on that a little bit already in This is not an atom.

    But, please, tell me now, are they in two places?

    No, not even technically. Linguistically then? Also, no. That statement is likely the result of mixing-up or lack for a better way of providing both metaphorical and physical descriptions of what is going on โ€“ it also reveals the still-present, outdated notion of what photons and electrons were supposed to be. If electrons were in two places at once, you still imagine them being small, little pellets, two copies of which fly through both slits, somehow interfering with each other. And thatโ€™s just not so.

    The correct expression is that electrons and photons and the likes donโ€™t have a definite location: their existence is simply spread out in space according to the wave function, the time-evolution of which in turn obeys the Schrรถdinger equation. Again, do give This is not an atom a read where this is explained in more detail.

    Pretty mind-bending stuff, right? Good. Welcome to the club. Great minds before you have had to take their time to wrap their heads around the double-slit experiment. Now youโ€™re one of them.


    Featured image by Free-Photos

    Sunlight diffraction pattern by Aleksandr Berdnikov under CC BY-SA 4.0

    Single electron build-up series. Results of a double-slit-experiment performed by Dr. Tonomura showing the build-up of an interference pattern of single electrons. Numbers of electrons are 11 (a), 200 (b), 6000 (c), 40000 (d), 140000 (e). By Belsazar under CC BY-SA 3.0.

  • The Collatz Conjecture

    The Collatz Conjecture


    This Conjecture is probably one of the easiest to understand which hasnโ€™t yet been proven in the history of mathematics. The beauty of this one is that a student in the last forms of primary school might very well be able to do the calculations while, thus far, not even the greatest mathematical minds have been able to prove if and why the Collatz Conjecture is true or not. The great, late Hungarian mathematician Paul Erdล‘s has been quoted as saying: โ€˜Mathematics may not be ready for such problems.โ€™[1]

    The rules

    There is some controversy over whether the prolific German mathematician Lothar Collatz was actually the first to come up with the idea in 1937, two years after his receiving his doctorate. It is also known as the โ€˜3n + 1 problemโ€™, the Ulam conjecture, Kakutaniโ€™s Problem, the Thwaites Conjecture, Hasseโ€™s Algorithm or the Syracuse Problem. If you were under the impression most of these refer to other, actual people then you are correct.

    As the rules of the Conjecture are so simple, it is likely many people have had the same idea independently of one another.

    Here are the rules:

    1. Take any positive, whole number โ€“ a positive integer, as itโ€™s called.
    2. Do either of the following:
      • if the number is even, divide by 2;
      • if the number is odd, multiply by 3, add 1.
    3. Take the result and do either of the following:
      • if the result is 1, stop;
      • if the result is not 1, go back and do step 2 again but this time using the result to do either of the two operations, and so on.

    The Collatz Conjecture goes as follows: no matter which positive integer you start from, irrespective of the number of steps, you will always get 1 as final outcome.

    (Note that if you would continue to do the steps with 1, you would simply cycle back to 1 in just three steps until the end of times. I mean, thatโ€™s just boring and silly. So, stop at 1.)

    Example

    Letโ€™s try it out. Letโ€™s start with 10.
    10 is even; divided by 2 equals 5.
    5 is uneven; multiplied by 3 plus 1 equals 16.
    16 is even; divided by 2 equals 8.
    8 is even; divided by 2 equals 4.
    4 is even; divided by 2 equals 2.
    2 is even; divided by 2 equals 1. We’re there!

    This took us 6 steps. You can try a few numbers yourself if youโ€™d like. Well, that is, have your computer, mobile phone or tablet do the boring work, which you can do here.

    Visualisation

    As you probably saw via the link above, we can also visualise the steps produced by the Collatz algorithm. We will quickly show a few alternatives before moving on to the most famous one, the Edmund Harriss visualisation (for which we wrote a JavaScript applet, yay!).

    Suppose, we would plot the progression of the results of our example. We started with 10. As you can see, the values fluctuate a bit before descending to 1:

    Have a look at the next one. We started at 8000. The numerical progression is then 8000 โ†’ 4000 โ†’ 2000 โ†’ 1000 โ†’ 500 โ†’ 250 โ†’ 125 โ†’ 376 โ†’ 188 โ†’ 94 โ†’ 47 โ†’ 142 โ†’ 71 โ†’ 214 โ†’ 107 โ†’ 322 โ†’ 161 โ†’ 484 โ†’ 242 โ†’ 121 โ†’ 364 โ†’ 182 โ†’ 91 โ†’ 274 โ†’ 137 โ†’ 412 โ†’ 206 โ†’ 103 โ†’ 310 โ†’ 155 โ†’ 466 โ†’ 233 โ†’ 700 โ†’ 350 โ†’ 175 โ†’ 526 โ†’ 263 โ†’ 790 โ†’ 395 โ†’ 1186 โ†’ 593 โ†’ 1780 โ†’ 890 โ†’ 445 โ†’ 1336 โ†’ 668 โ†’ 334 โ†’ 167 โ†’ 502 โ†’ 251 โ†’ 754 โ†’ 377 โ†’ 1132 โ†’ 566 โ†’ 283 โ†’ 850 โ†’ 425 โ†’ 1276 โ†’ 638 โ†’ 319 โ†’ 958 โ†’ 479 โ†’ 1438 โ†’ 719 โ†’ 2158 โ†’ 1079 โ†’ 3238 โ†’ 1619 โ†’ 4858 โ†’ 2429 โ†’ 7288 โ†’ 3644 โ†’ 1822 โ†’ 911 โ†’ 2734 โ†’ 1367 โ†’ 4102 โ†’ 2051 โ†’ 6154 โ†’ 3077 โ†’ 9232 โ†’ 4616 โ†’ 2308 โ†’ 1154 โ†’ 577 โ†’ 1732 โ†’ 866 โ†’ 433 โ†’ 1300 โ†’ 650 โ†’ 325 โ†’ 976 โ†’ 488 โ†’ 244 โ†’ 122 โ†’ 61 โ†’ 184 โ†’ 92 โ†’ 46 โ†’ 23 โ†’ 70 โ†’ 35 โ†’ 106 โ†’ 53 โ†’ 160 โ†’ 80 โ†’ 40 โ†’ 20 โ†’ 10 โ†’ 5 โ†’ 16 โ†’ 8 โ†’ 4 โ†’ 2 โ†’ 1.

    Thatโ€™s a whole lot of numbers before the algorithm leads to 1. Hundred and fourteen steps, to be precise. This is what it looks like:

    As you can see, the whole plot fluctuates quite a bit. Itโ€™s a bit of a mess, really. Thereโ€™s no distinct pattern other than it eventually converging to 1. As the values may become quite large very quickly, letโ€™s plot the same graph in a semi-log grid. The y-axis is logarithmic, the x-axis remains linear.

    Just to humour ourselves, let’s reverse the step order, so that the plot is mirrored and converges to the value 1 in the bottom-left corner:

    As weโ€™re not sure how many steps it might take before a sequence of numbers ends with 1, letโ€™s also change the x-axis to a log scale. We get this:

    With this one, we can probably visualise a whole bunch of sequences! Lastly, letโ€™s now plot a series of sequences! That is, firstly, we let our JavaScript applet calculate the sequence starting at 10000. Then we let it calculate the sequence starting at 9999. And so on, downwards, until it reaches 4. And then have it all plotted, all at once! To prevent it from becoming too dense, as several sequences will overlap each other, we add a little transparency to each plot. If the same โ€˜pathโ€™ has been taken, that path will appear โ€˜darkerโ€™. This is the result:

    This almost becomes some form of art. If you would frame this plot, without the titles, axes, and scales โ€“ just the plot โ€“ you could have a piece of geometrical art bearing the title โ€˜The Collatz Conjectureโ€™ or something like that. I might do that, actually.

    With this applet you can generate your own โ€˜artโ€™ like the one above.

    The Edmund Harriss visualisation

    Edmund Harriss, a mathematician working at the University of Arkansas, came up with a beautiful visualisation of progressions of Collatz sequences, which was featured on the mathematical YouTube channel Numberphile. We highly recommend following their channel.

    A screenshot of Numberphileโ€™s video showing a partly hand-made version of Edmund Harrissโ€™s visualisation

    The rules of visualisation are simple. While iterating through the Collatz rules, the algorithm goes as follows. If the current step is twice the value of the next step, rotate a fixed amount clockwise, otherwise rotate half of that fixed amount anticlockwise (and, again, stop at value 1). The result is a bundle of threads which looks like some sort of organic entity โ€“ messy and seemingly randomised within certain constraints, just like nature.

    We wrote a JavaScript applet to try and produce an approximation of his visualisation. We used it to produce the featured image at the top of this page. You can try it here yourself.

    Having visualised in several ways the kind of disorderly fashion in which the Collatz sequences progress, thus far, it may not seem surprising it has proven to be hard to crack the code. Well, the underlying mathematical code that is, not the JavaScript code.

    [1] Guy, Richard K. (2004). “E17: Permutation Sequences”. Unsolved problems in number theory (3rd ed.). Springer-Verlag. pp. 336โ€“7. ISBN 0-387-20860-7. Zbl 1058.11001.

  • This is not an atom

    This is not an atom


    Today, many people know that all things around us โ€“ the chair we sit on, the screen we look at โ€“ are a composition of all sorts of different molecules and that they are in turn composed of all sorts of different atoms. Indeed, ancient Greek philosophers such as Democritus hypothesized matter consists of tiny, physically indivisible entities, which they then named atoms.

    However, we also know that the Greeks werenโ€™t entirely correct: the atom itself is composed of electrons, protons, and neutrons. We also know that the latter two are composed of even smaller things โ€“ quarks and gluons.

    What not many people know, however, is that this classical picture:

    A out-dated image of an atom. Several tiny balls fly in fixed orbits

    is absolutely not what an atom is!

    Old ideas

    If you hung out in the wrong street corners, you might have been told that electrons whizz around the nucleus like tiny planets around the Sun or tiny moons around a planet. If that were the case then you have been lied to.

    If you hung out in yet other unsavoury street corners, you might have been told that quantum mechanics is something magical, spiritual, and the doorway to a deeper understanding of love, consciousness, and healing. Telepathy, even. Again, you have been lied to.

    Admittedly, the famous physicist Richard Feynman is often quoted as saying that nobody understands quantum mechanics. In a specific way, thatโ€™s true. Particles donโ€™t behave like everyday objects and that is a strange fact. Furthermore, the mathematical descriptions of particles tell us what they do but not what they are. We know all the equations but we donโ€™t know what they mean โ€“ as opposed to knowing the meaning of the words โ€˜microscopically tiny ballโ€™.

    However, this doesnโ€™t mean we shouldnโ€™t make an effort to making particles predictable, useful, and less mysterious and esoteric. It doesnโ€™t mean we canโ€™t harness the power of a good theory of quantum behaviour.

    In fact, thatโ€™s exactly what weโ€™ve been doing rather successfully since quantum mechanics took shape in the 1920s. Hence, the existence of your mobile phone, computers, cameras, and self-checkout in the supermarket.

    So, letโ€™s slice off the fat and cut to the chase.

    Classical mechanics versus quantum mechanics

    In high school, we were taught Newtonian mechanics. We were told that the world is reigned by Newtonโ€™s laws, the most powerful of them being the second: force equals mass times acceleration,

    $F = ma.$

    We were taught that when an object is moving, it moves according to Newtonโ€™s second law.  The beauty of his mechanics was that physicists and engineers were now able to predict the future (and retrodict its past) of a sliding block, for instance, based on just a few known initial conditions.

    Inclined plane problems are the staple in physics class for senior high school students

    More generally speaking, and in physics jargon, Newtonian mechanics is capable of describing the state of a system over time with mathematically infinite precision based on a sufficient set of initial conditions. We call this a deterministic theory as itโ€™s possible to determine past and future of the state of a system. Thanks to this property even space vessels such as the Apollo Lunar Module and Mars Rover Curiosity were able to arrive successfully at their extraterrestrial destinations.

    In quantum mechanics, we study the behaviour of subatomic stuff, such as electrons and quarks. After many twists and turns throughout history, it turns out we canโ€™t actually determine the past and future of, say, an electron โ€“ not as we could for blocks and balls in Newtonian mechanics. Why not? Because itโ€™s simply not a tiny block or ball. Itโ€™s not even a particle in that sense (assuming a particle is like a tiny ball)! Itโ€™s a wave function, a mathematical expression describing all the possible states of a โ€˜particleโ€™. Quite a different beast.

    Possible states? Yes, in quantum mechanics, things arenโ€™t so deterministic. Turns out that to describe the state of an electron, for example, Newtonโ€™s second law doesnโ€™t apply. Itโ€™s fundamentally impossible to predict or retrodict where an electron will be at any given time, for instance. Or how fast itโ€™s moving at a particular point in time. The best we can do is calculate the probability itโ€™ll be here or there or whizzing at this or that velocity. In other words, the state of an electron can only be described in terms of probabilities.

    It also turns out that the probabilities of this set of possible states may change over time. Luckily, like in Newtonian mechanics, thereโ€™s an equation for that. In quantum mechanics, the analogue of Newtonโ€™s second law is called the Schrรถdinger equation. It tells you how a wave function, i.e., the set of all possible states of a particle, changes over time. In its most compact form(beginfootnote)Although, technically, using Newtonโ€™s notation instead of Leibnizโ€™s, an even more compact form is $i \hbar \dot{\Psi} = \hat{H} \Psi.$(endfootnote) it goes like this:

    $i\hbar \dfrac{\partial}{\partial t}\Psi = \hat{H}\Psi.$

    No need to understand all the symbols but here you can see that also in quantum mechanics thereโ€™s a beautiful equation at its centre, and itโ€™s this one(beginfootnote)There are other ways to calculate the time-evolution of the wave function, of course, such as in Heisenbergโ€™s matrix mechanics, Feynmanโ€™s path integral formulation, and Diracโ€™s formulation for matrix mechanics and the Schrรถdinger equation combined. However, this one is invariably taught at undergraduate level.(endfootnote). It tells you the evolution of $\Psi$, the symbol for the wave function.

    Deterministic theories, such as Newtonian mechanics, are called โ€˜classicalโ€™ as opposed to quantum mechanics, dealing with probabilistic wave functions(beginfootnote)Note that the Schrรถdinger equation is deterministic. Itโ€™s the wave function itself that yields probabilities, or, if youโ€™re a stickler for accuracy like me, itโ€™s the wave functionโ€™s norm squared that yields probability densities (by integrating the norm squared over a volume, area or distance).(endfootnote).

    Wave functions

    As weโ€™ve learnt in the previous section, a particle is not a particle. Granted, we still talk about a โ€˜particleโ€™ but thatโ€™s only for lack of a better term. Itโ€™s an artefact of humankindโ€™s limited understanding of the Universe back in the day. The idea of a particle simply fits among the things we already know. We can picture a little ball because we grew up playing with little balls. Or marbles, or whatever. Admittedly, sometimes particles do look like particles, which weโ€™ll discuss in the last section.

    Nevertheless, experiments from the early 20th century proved that tiny balls were definitely the wrong idea. Therefore, nowadays, our best descriptions of โ€˜particlesโ€™ are indeed wave functions, mathematical expressions. The fundamental question is whether the wave function is the particle or just a mathematical representation of it. This question hasnโ€™t been answered yet, however, the personal opinion of the author of this post is that after about a century of the highly successful theory of quantum mechanics, itโ€™s maybe time to start regarding the wave function as the thing that is a โ€˜particleโ€™.

    Just to illustrate the difference between a particle and wave, have a look at this point-like particle. The horizontal axis is the x-coordinate in space and the vertical axis is the y-coordinate in space.

    Now tell me, where in space is the particle located? You probably got the answer straight away. Itโ€™s at coordinate (2,3). Good.

    Now, have a look at this (two-dimensional) wave.

    So, tell me, where in space is the wave located? You may find it harder to pinpoint the wave to a specific set of coordinates. Thatโ€™s because itโ€™s in several places at once. It doesnโ€™t have a specific position. In physics speak, this is called a superposition.

    This is also the case for an electron (or any other โ€˜particleโ€™). Itโ€™s in a superposition, and not just in terms of its position: itโ€™s also in a superposition in terms of its energy, momentum, and a few other properties. In other words, itโ€™s in all places at once at several energy levels at once, whizzing at all kinds of velocities at once.

    Note, however, that the wave function is absolutely not the same as a simple sine wave in normal space which was merely displayed here for reasons of clarity(beginfootnote)And, indeed, those who read a previous post on light refraction in glass know that โ€˜particlesโ€™ are oscillations in their respective three-dimensional quantum fields in quantum field theory. The wave function plays a central role in this highly successful theory. Secondly, the wave function is a so-called complex function and therefore exists in so-called complex space $\mathbb{C}$, not in the โ€˜regularโ€™ number space $\mathbb{R}$, we all grew up with. Lastly, and at the cutting edge of our scientific knowledge, there is actually only one wave function, the wave function of the Universe. Every field and particle in it are mere parts of that wave function, which can be thought of as small, individual wave functions to keep it manageable.(endfootnote).

    Maybe now you can appreciate how revolutionary quantum mechanics truly is compared to the simple mechanics of the blocks and tiny balls of everyday life.

    Picture of an atom

    So, weโ€™ve arrived at the correct picture of an atom. We already learnt that an electron isnโ€™t a point-like particle, itโ€™s a wave function. What do wave functions look like then? Well, they are cloud-like but not clouds, smeared out in space, yet both size- and location-less. They donโ€™t have a specific position, they donโ€™t have a specific momentum, they donโ€™t have a specific energy value. They are in a superposition of all these possible states. The probabilities of these states may oscillate over time as dictated by the Schrรถdinger equation.

    That doesnโ€™t help much, does it?

    Well, fear not. Dillon Berger, a PhD student of Theoretical Particle Physics at UC Irvine, made a beautiful animation of a cross-section of a hydrogen atom using the Schrรถdinger equation. The contours represent the wave function of the electron. The colours denote the probability of the electron being in that particular state. Note, there is only one electron in a hydrogen atom. So, yes. Itโ€™s almost everywhere at the same time, while also oscillating over time. (This animation has time slowed down by a thousand trillion. The nucleus, a proton, is too small, so itโ€™s invisible.)

    That whole tiny balls or planets revolving around the nucleus analogy? Flush it out of your system. For good.

    Unless weโ€™re looking

    Now, hold on, you might say. Why is it then that professional physicists still talk of particles? And what did you mean, when, earlier, you said theyโ€™re size-less? How then do you explain the fact that in scientific tables we saw in high school, actual, physical sizes of particles are listed? And how do you explain this classical picture of the readout of a cloud chamber demonstrating the existence of a subatomic particle? That trajectory certainly looks like it was created by a point-like particle and not at all a wave.

    Source: Anderson, Carl D. (1933). “The Positive Electron”. Physical Review 43 (6): 491โ€“494

    Youโ€™re quite right to doubt the whole story about wave functions in the face of these empirical findings. Youโ€™ve also arrived at a mystery that is at the heart of quantum mechanics, worthy of a Nobel Prize, which, of course, has a name: the measurement problem.

    Turns out, particles are indeed wave functions but only if weโ€™re not looking. As soon as we do measurements, trying to gauge their position, for example, we wonโ€™t find them at all places at once, like a wave. Instead, we will find them at one particular location โ€“ just as we would expect from an actual particle!

    This is what the famed double-slit experiment demonstrated. In another post, we discussed this.

    This is why, to this day, you might have heard of the โ€˜wave-particle dualityโ€™ of the subatomic world. The wave function is the most complete description of a particle. As soon as we do measurements, we see only a sliver of its original wave function, a mere shard of the set of all possible states.

    Therefore, Dillon Bergerโ€™s animation shows a hydrogen atom when left completely alone. This is its fundamental state of being: its electron being a wave function in superposition, oscillating over time according to the Schrรถdinger equation. And as soon as it interacts with its environment, only a metaphorical slice of its full existence will show (a slice corresponding to the disguise of an actual particle).

    How this happens or what actually happens when we do measurements is still up for debate. We will most certainly dive deeper into a variety of views on how to tackle this phenomenon in another post. Expect a post on the Copenhagen interpretation of quantum mechanics, the Many-Worlds interpretation, and others soon.

    Now you have a better mental picture of an atom, at least. Probably.

  • Proof that the square root of 2 is irrational

    Proof that the square root of 2 is irrational


    While itโ€™s one of the most well-known and well-trodden proofs among proofs, the irrationality of $\sqrt{2}$ shouldnโ€™t be lacking on a blog about mathematics and physics. So, here it goes.

    What is irrationality?

    For those who arenโ€™t too familiar with mathematical jargon, letโ€™s first discuss what it means to be irrational. Obviously, weโ€™re not talking about the psychological attribute but the mathematical one.

    You might remember primary school when you had to learn about fractions such as

    \begin{equation} 1 = \frac{4}{12} + \frac{2}{3}. \end{equation}

    A practical application of a fraction is when you were reading a recipe for a delicious dish with a certain ratio of water and rice, which, even if you might not be aware of it all the time, can be written as a fraction, representing the ratio between water and rice. In fact, a fraction is a ratio.

    For instance, in order to cook the perfect, fluffy rice without the need to pour off excess water when the rice is cooked, the ratio is that for 1 cup of rice, you add 1.5 cups of water(beginfootnote)Rinse the rice thoroughly to remove the starch and dust for a nice fluffy texture. Add water by 1.5 times the used volume of rice. Add salt if you must. Bring the water to the boil as quickly as possible. Bring down the heat but keep the water bubbling softly. Give it one good stir. Put the lid on and donโ€™t remove it for eight minutes. Donโ€™t look inside; the water needs to stay in the pan. After eight minutes, shut off the heat and let it rest for another eight minutes. Still, donโ€™t look. Keep the lid on the whole time. Thatโ€™s it.(endfootnote) So, the fraction is $\frac{1}{1.5}$.

    Of course, itโ€™s conventional to write a fraction using whole numbers (integers) only, so, $\frac{1}{1.5} = \frac{2}{3}$. Just multiply the numerator and the denominator by two. In other words, for 2 cups of rice, add 3 cups of water.

    If we use our calculator, we get $\frac{2}{3} = 0.666\dots$ There is no end to this number, but the number can be perfectly written down as a ratio: $\frac{2}{3}$.

    Of course, $\frac{2}{3}$ is the same as $\frac{4}{6}$, or $\frac{10}{15}$, or $\frac{200}{300}$, since, and this is crucial, all the other fractions (ratios) are simply multiples of our original fraction: they can all be simplified to their โ€˜simplestโ€™ form, $\frac{2}{3}$. A slightly more technical way of saying this is that the fraction $\frac{2}{3}$ is the form in the lowest terms of the fraction $\frac{200}{300}$. Itโ€™s very important to remember this.

    Now weโ€™ve arrived at what irrational numbers are.

    Premise 1. A number is irrational when it cannot be written as a ratio in lowest terms.

    Two of the more well-known examples of irrational numbers are $\pi$ and $\sqrt{2}$. If we use our calculator, we can see how there seems to be no numerical repetition in them. This is a quality that irrational numbers possess.

    Babylonian tablet clay tablet showing the root of 2 (credits below)

    Proof by contradiction

    So, how do we prove that $\sqrt{2}$ is irrational, i.e. it cannot be written as a ratio? We do this by contradiction: if the opposite of a statement is demonstrably false (and there are really only two options), then the statement itself must be true. In a previous article, we used the same strategy to prove that โ€˜minus minus is plusโ€™.

    We will do that here, too.

    Even and odd

    Premise 2

    We will also use the fact that some number multiplied by 2 equals an even number. Take any number, odd or even, multiply that by 2, and you will get an even number. This isnโ€™t rocket science, really, as a characteristic of an even number is that itโ€™s divisible by 2. In our proof, we will use the symbol $k$ for โ€˜some number, any number, odd or evenโ€™ being multiplied by 2.

    Premise 3

    If you take the square of an odd number, the result is always odd. If you take the square of an even number, the result is always even. Conversely, if you take the root of an odd number, the result is always odd. The same idea goes for even numbers. Check in your head to see if thatโ€™s true (it is). We will provide a proof for that in another post.

    Okay, ready? Letโ€™s go.

    Proof that the square root of 2 is irrational

    Anti-Premise 1. Suppose, by contradiction, that $\sqrt{2}$ can be written as some ratio in lowest terms: some (whole) number $a$ divided by some other (whole) number $b$ in lowest terms.

    In other words, suppose

    \begin{equation} \sqrt{2} = \frac{a}{b}. \end{equation}

    To make life a little bit easier, we get rid of the square root by squaring both sides of the equation:

    \begin{equation} 2 = \frac{a^2}{b^2}. \end{equation}

    If we rearrange this, we get

    \begin{equation} a^2 = 2b^2. \end{equation}

    Now, we see that $a^2$ is an even number as $b^2$ โ€“ whatever that number is โ€“ is multiplied by 2. It also means that $a$ cannot be an odd number โ€“ itโ€™s even. Remember premise 3?

    Conclusion 1: $a$ cannot be odd โ€“ itโ€™s even.

    We can then also state that $a = 2k$, where $k$ is some number, any number, odd or even. If we substitute that into equation (4), we get

    \begin{equation} (2k)^2 = 2b^2. \end{equation}

    If we rearrange that, we get

    \begin{equation} b^2 = \frac{(2k)^2}{2}. \end{equation}

    If we simplify this in one extra step, we get

    \begin{equation} b^2 = \frac{4k^2}{2} = 2k^2. \end{equation}

    This means that irrespective of what number $k^2$ is, because itโ€™s multiplied by 2, the result is an even number. In other words, $b^2$ is an even number, which also means that $b$ is an even number.

    Conclusion 2. $b$ cannot be an odd number โ€“ itโ€™s also an even number.

    Looking at conclusions 1 and 2, we arrive at

    Conclusion 3: $\frac{a}{b}$ is not a ratio in the lowest whole numbers as $a$ and $b$ can still be divided by 2.

    This is a contradiction. The fraction $\frac{a}{b}$ cannot both be the lowest fraction and not be the lowest fraction. Conclusion 3 contradicts Anti-Premise 1. Therefore, there is no fraction $\frac{a}{b}$ in lowest terms that exists that can be equal to $\sqrt{2}$. Hence, the original statement Premise 1 is true.


    Credentials of the Babylonian tablet clay tablet showing the root of 2: Photograph by Bill Casselman under CC BY-SA 3.0, and the Yale Babylonian Collection as the original holder of the tablet. A black and white rendition of Casselmanโ€™s own photograph of the Yale Babylonian Collection‘s Tablet YBC 7289 (c. 1800โ€“1600 BCE), showing a Babylonian approximation to the square root of 2 (1 24 51 10 w: sexagesimal) in the context of Pythagoras’ Theorem for an isosceles triangle. The tablet also gives an example where one side of the square is 30, and the resulting diagonal is 42 25 35 or 42.4263888…(30 x square root of 2).


  • The meaning of E=mcยฒ

    The meaning of E=mcยฒ


    Probably the most famous equation on this planet is $E = mc^2$. Energy equals mass times the speed of light squared(beginfootnote)Usually, in mathematics, we leave out the multiplication sign ($\times$).(endfootnote). Usually, the formula is associated with Albert Einstein. This relationship between, energy, mass, and the speed of light, this equation, has a name: the mass-energy equivalence. Perhaps youโ€™ve read or heard people explain that โ€˜Einstein taught usโ€™, that mass is a form of energy, mass is frozen energy, mass can be converted into energy, and that, in the end, all matter is essentially pure energy.

    The equation is, however, definitely not about all of that. At least, not in this Universe. Here, we will discuss the actual meaning of it. We think itโ€™s time for disposing of some of the unnecessary obscurantism accompanying many popular explanations. We think the involved mathematicians and physicists of yore deserve better.

    Albert Einstein (right) with Dutch physicist Paul Ehrenfest (left) and Ehrenfest’s son in Ehrenfest’s home in Leiden, The Netherlands.

    Standing on the shoulders of giants

    Einstein wasnโ€™t the first to write down this very relationship between mass and energy. There were many others before him who, one way or the other, explored the connection between mass, energy, and velocity. However, most hypothesised that a mechanical mass increase was exclusively due to interactions with electromagnetic fields. They called it electromagnetic self-energy of some kind, giving rise to a form of electromagnetic mass (Miller, 1981; Okun, 1989). Einstein then showed that there was no need for such a concept and was the first to derive the relation correctly.

    However, over the course of a few years, he published several derivations of the equation, none of which were literally written down as $E = mc^2$. You might not recognise them if you saw them. Furthermore, the versions he did write down, werenโ€™t universally true.

    Lastly, the famous equation isnโ€™t the complete version. Usually, people only know the snazzy edition fitting on baseball caps. The full equation is valid in a more universal way and sometimes referred to by contemporary physicists as the โ€˜correct versionโ€™(beginfootnote)cf. https://youtu.be/mkiCPMjpysc(endfootnote). However, this one wasnโ€™t first formulated by Einstein but by Paul Dirac (Eisberg & Resnick, 1974; Miller, 1981). We will get back to that.

    Among others, these incredible minds have all derived and used some version of the famous equation before Einstein.

    Energy is a mathematical idea

    We should remind ourselves that energy isnโ€™t any โ€˜thingโ€™. As we mentioned in a previous article, it isnโ€™t some invisible, immaterial, fluid-like โ€˜essenceโ€™ which everything is made of, within and behind the faรงade of the tangible world. Fork, no.

    It has always been just a number, an accounting tool, an important and practical, mathematical concept, first proposed by the 17th century, German scholar Gottfried Wilhelm von Leibniz. How is it a mathematical concept? Itโ€™s the number you get when you multiply an objectโ€™s mass with its velocity squared(beginfootnote)Which was later calibrated to 1/2 times the mass times velocity squared.(endfootnote). Itโ€™s a pragmatic way of keeping the books on these two things.

    Suppose, we have a billiard table with three billiard balls. We ignore any sort of friction. Imagine this isolated system of balls changing internally: the balls constantly collide and bounce off of the edge of the table. He assumed that what doesnโ€™t change, is their mass. What does change, are their velocities. Leibniz then noticed that if you multiply for each ball its mass with its velocity squared, and summed all three products, that sum remained constant, irrespective of how the balls were bouncing in which direction, and how fast, at each point in time (until you changed something to the system by introducing a whack by a cue stick, for example).

    A sketch of a carom billiards table. First panel: three balls on the billiards table have different velocities. Second panel: the three balls have bounced and moved and have, again, different velocities. The sum of the products of the mass of the balls with its velocity squared, of all balls, is constant, however.

    This was the early formulation of what we now fanciful call the law of conservation of energy(beginfootnote)He called this product quantity vis viva, so, not even energy yet. Perhaps he was being poetic. Furthermore, Leibniz also had some fierce competition: his rival Newton had come up with a different quantity, a different way of keeping the books. He stated that the sum of the products of mass and velocity โ€” just velocity, not velocity squared โ€” remained constant. This is the law of conservation of momentum. Later, it was understood that the two laws were complimentary, not contradictory.(endfootnote), which, by the way, is true in a small enough patch of the Universe, such as Earth or the solar system, but in the context of the entire observable Universe, for instance, energy is not conserved. So, indeed, the law of conservation of energy is fundamental enough for us, Earthlings, and our physics experiments, however, contrary to what people usually think, in the grander scheme of the Universe, itโ€™s not(beginfootnote)Courtesy of the genius Emmy Noether (Noetherโ€™s theorem) and Albert Einstein (general relativity).(endfootnote).

    For the purpose of this post, however, this is irrelevant. What is important to note, is that since then, with their propensity to invent intricate systems of categorisation, humans have made distinctions between various forms of energy. Think of potential energy (somethingโ€™s high up and can fall down or wound up and unwind rapidly), thermal energy (somethingโ€™s hot), chemical energy (somethingโ€™s โ€˜chargedโ€™), and kinetic energy (something moves). All of these, including their mother-concept โ€˜energyโ€™, are human constructs. Thinking of any of these nouns as referring to physically separate, physical, tangible things is, ironically, a category mistake.

    Energy isn’t an ephemeral and/or ethereal substance. It’s a mathematical measure for the product of mass or momentum, and speed. In this hastily taken photograph by an eyewitness, it’s not beams of pure energy that you’re seeing, even though this does appeal more to our imagination of what โ€˜pure energyโ€™ is supposed to be. They are particle rays (orange-coloured bundles of radially polarised protons), emitted by portable particle accelerators on their backs, aimed at a ghost. The Ghostbusters, as they call themselves, stated that ghosts are negatively charged energy in the form of slime-like ectoplasm. So, even here โ€” fictional or not โ€” when thereโ€™s something strange in your neighbourhood, it’s never ‘pure energy’. (ยฉ Sony Pictures Home Entertainment)

    It’s not all about that mass

    Sometimes, people misinterpreted Albert Einstein. In the old days, again, being the talented labellers that they are, humans split up the term โ€˜massโ€™ into rest mass and relativistic mass. Rest mass is the mass when the object is at rest. Relativistic mass is the mass when the object is in motion. If then the object would start to move faster and faster, then this particular mass would become larger and larger, because, you know, thatโ€™s what he said.

    Well, no. He wrote in his third 1905 paper (1905a, p. 920), Zur Elektrodynamik bewegter Kรถrper, an equation for the kinetic energy of an electron, which went as follows:

    While it may not look like it, you could say that this was his first expression for the relationship between (kinetic) energy, mass, and speed of light squared(beginfootnote)Incidentally, Max Abraham had published Walter Kaufmanโ€™s work just before Einstein, showing the same equation for kinetic energy. Einstein probably wasnโ€™t aware of this (Miller, 1981).(endfootnote). It does require a little translation, but itโ€™s easy. Ignore the part in the middle, focus on the letter $W$, and the part behind the last equals sign. You should know that in modern notation, kinetic energy $W=E_k$, rest mass $mu=m_0$, and speed of light $V=c$. So, what it says is:

    \begin{equation} E_k = \frac{m_0c^2}{\sqrt{1-\dfrac{v^2}{c^2}}} – m_0c^2. \end{equation}

    So, this is slowly starting to resemble our familiar $E=mc^2$. It doesnโ€™t state, however that mass increases. It only says that when the speed of the object $v$ approaches the speed of light $c$, the result of this whole equation is infinity โ€“ infinite energy.

    As the standing interpretation is that mass equals energy (mass-energy equivalence), people nevertheless concluded that if the energy of a moving object becomes infinite at the speed of light, that an objectโ€™s mass becomes infinite, or relativistic mass, to be precise (in the minds of the old folk).

    This, however, is something of the past โ€“ well, technically. As soon as 1940, the great Lev Landau and Evgeny Lifshitz ignored the distinction between rest mass and relativistic mass in their book The Classical Theory of Fields. The legendary John A. Wheeler and Edwin F. Taylor also brought an up-to-date Spacetime Physics to the reading table. Unfortunately, many textbooks today still mention archaic notions, terms, and notation.

    Contemporary professional physicists donโ€™t speak of relativistic mass anymore. In special relativity, Einstein showed that observations and measurements depend on oneโ€™s frame of reference. By definition, there are at least a couple of things that do not depend on the motion of observers. Besides the spacetime interval, the laws of physics, and the speed of light, this turns out to be rest mass.

    An objectโ€™s rest mass is invariant, i.e. it doesnโ€™t vary or change, regardless of the motion of the observer relative to the object (Taylor & Wheeler, 1992, p. 211). And, as you can see in Einsteinโ€™s equation for the kinetic energy, only rest mass is used. No mentioning of relativistic mass whatsoever. In fact, Einstein himself wrote (as cited in Okun, 1989, p. 32):

    โ€˜It is not good to introduce the concept of mass $M = m/(\sqrt{1-v^2/c^2})$ of a body for which no clear definition can be given. It is better to introduce no other mass concept than the โ€˜rest massโ€™ $m$. Instead of introducing $M$ it is better to mention the expression for momentum and energy of a body in motion.โ€™

    Of course, $M$ is what humans would later call โ€˜relativistic massโ€™, which means that they did it anyway, against Einsteinโ€™s wishes.

    Today, however โ€“ well, at least since 1940 โ€“ professional physicists speak only of mass. We tossed out relativistic mass as, with Einstein, itโ€™s โ€˜not goodโ€™. Furthermore, the adjective โ€˜restโ€™ in rest mass is redundant. Mass is about an object at rest. If itโ€™s in motion, we speak of the product of some proportion of mass and velocity: either momentum or energy. An objectโ€™s mass does not grow by its motion.

    Resistance is crucial 

    So, what is mass then? Well, just to be clear, mass isn’t weight: the same amount of mass has different weight on different planets. We were taught this in high school. A spring scale measures weight, not mass. A balance with calibrated counter-weights โ€“ *cough* masses โ€“ is your best option during interplanetary travels. These masses should have been calibrated to the definition of a kilogram according to the International Bureau of Weights and Measures.

    Mass isnโ€™t matter either. Mass and matter are different categories. The first is a property, the second is a โ€˜thingโ€™. Elementary โ€˜particlesโ€™, things(beginfootnote)We use quotations marks in โ€˜particlesโ€™ because, while itโ€™s easier to use that term, we acknowledge itโ€™s actually quantum field oscillators we should be talking about or, even prettier, wave functions.(endfootnote), such as the electron, have a certain amount of mass. An electron gains its mass through interacting with the Higgs field, the existence of which was proved in 2012 at CERN. And the amount of elementary particles does correlate with the amount of mass. However, the mass of an object isn’t defined by merely the amount of elementary particles: it’s also the motion of gluons inside protons, the motion of electrons, atoms, molecules โ€“ the kinetic, thermal, and chemical energy contained within the (resting) object.

    Einstein wrote this in his fourth paper of 1905, Ist die Trรคgheit eines Kรถrpers von seinem Energieinhalt abhรคngig? (1905b, p. 641):

    If a body releases the energy L in the form of radiation, its mass decreases by $L/V^2$,

    where, in modern notation, $L=E$, and $V=c$. Also, the energy heโ€™s talking about is not the kinetic energy (of a body in motion) but the internal energy (of a body at rest), such as thermal energy. And yes, this does mean that mass increases or decreases depending on the objectโ€™s internal energy.

    Two structurally identical balls of steel have different mass if one ball is hotter (more mass) than the other due to their diverging thermal energy content. Two structurally identical mobile phones have different mass if one is charged (more mass) and the other is out of juice (electrochemical energy). Note, we are talking about the whole object being at rest in its reference frame.

    As soon as either object starts radiating light or heat, they lose mass. The fraction $E/c^2$, however, is very small because the speed of light is very high. And so, the extra mass gained or lost is so small that this may be a reason for people confusing mass with the amount of matter. Itโ€™s almost the same. Itโ€™s the amount of matter plus something more, its internal energy.

    Mass could be best described by the resistance to acceleration โ€“ a ratio between the force needed to accelerate it to the extent it’s accelerating. Mass is inertial mass, an object’s inertia (as Einstein put it in the title of his 1905a paper).

    The complete equation

    If you’re not into equations any longer, skip this section. If you want to know the real equation, letโ€™s go.

    Einstein published several derivations. One of the more familiar was the following for the total energy:

    \begin{equation} E_T = \frac{m_0c^2}{\sqrt{1-\dfrac{v^2}{c^2}}}. \end{equation}

    If you just to happen to be fluent in algebra, then you could see how we obtain $E=mc^2$. If not, not to worry. If an object is at rest, is has no speed, so $v=0$. If you would fill in that number in the equation, the denominator of the big fraction becomes the value 1. And anything divided by 1 equals exactly that same anything. So, that means that what you get is $E=m_0c^2$.

    Of course, since, nowadays, there is only one mass, which is $m$, since โ€˜restโ€™ is redundant, we should really leave out the subscript 0. This also means that the famous equation $E=mc^2$ is only applicable if the object isn’t moving in our (inertial) reference frame. Moreover, it’s not applicable to phenomena without mass either, such as a photon. Hence,

    $E=mc^2$ is not universally true. Only in an inertial reference frame, where the object isn’t moving, and only in the case of ‘particles’ with mass, does this equation hold, so, this equation isnโ€™t valid for photons and gluons.

    This following equation, however, does hold for massless as well as massive particles, and while Einstein laid the groundwork, the genius Paul Dirac was to write this down for the first time in 1928 (Eisberg & Resnick, 1974; Miller, 1981), albeit in a slightly more technical fashion than presented here. The following equation handles all objects, including light:

    \begin{equation} E^2 = m^2c^4 + p^2c^2. \end{equation}

    The letter $p$ is the momentum. Suppose, we want to calculate the energy of a photon. Since the photon has no mass ($m=0$), this equation becomes $E = pc$, which is, indeed, the correct relation between energy and a photon. If you would try to use $E = mc^2$ to calculate the energy of a photon, you would get a silly answer.

    So, $E = mc^2$ isn’t even a universal equation because it doesnโ€™t fly for massless โ€˜particlesโ€™: photons and gluons. The equation first written down by Paul Dirac does, however. And it still fits on a T-shirt. 

    Often, though, it’s written as

    \begin{equation} E^2 = (pc)^2 + (mc^2)^2, \end{equation}

    which makes it possibly even snazzier as it shows a beautiful Pythagorean relationship triangle.

    Paul Dirac

    The meaning of E = mcยฒ

    All well and good, but, technicalities aside, what does it mean?

    What it means is that energy is mass, proportioned by a factor of $c^2$.

    What it also means is that an objectโ€™s mass is a measure of its total intrinsic energy (potential, thermal, chemical, electrical, even kinetic, if parts inside the object have motion) proportioned by a factor of $1/c^2$.

    $E = mc^2$ should actually be written $E_0 = mc^2$ as itโ€™s about the energy of an object at rest and the subscript 0 usually denotes something at rest.

    However, it isnโ€™t the full equation.

    What it doesnโ€™t mean is that energy is matter, and, conversely, it doesnโ€™t also mean that matter is energy. Mass isnโ€™t matter. This is a category mistake.

    It also doesnโ€™t mean that mass can be converted into energy or vice versa. For one, mass cannot be converted as it isnโ€™t a โ€˜thingโ€™. Secondly, energy isnโ€™t a โ€˜thingโ€™ either. Mass is a property, a measurable property. Energy is also a property, a calculable property, which can be done by measuring mass.

    Imagine an object had the following properties: size, colour, hardness, and energy. Suppose, the equation would have said $E =$ hardness $\times c^2$. Perhaps itโ€™s more clear now that this doesnโ€™t mean that hardness gets converted into energy. What it means, is that you have a mathematical way of calculating one measure in terms of the other measure. The only thing thatโ€™s being converted here, is a number, a quantity.

    Matter is a different beast. Itโ€™s a clump of things: โ€˜particlesโ€™. An object is a clump of matter and matter interactions. As CERN show on a daily basis, matter in motion can be converted into a thousand other things in motion. If people insist on talking about things getting converted, then they could talk about converting particle A with motion $a$ and particle B with motion $b$ into particles C, D, E, F, G, $\dots$ with motions $c,d,e,f,g,\dots$

    So, next time someone thinks they should explain to you that $E = mc^2$ means that mass can be converted into energy or that no object can gain the speed of light because its mass would become infinite, you can just reply with, โ€˜Nah, mate, mass is an invariant property of an object, calculable through the complete equation, you know, $E^2 = (pc)^2 + (mc^2)^2$. Although you would have to solve for $m$ and merely use the pseudo-Euclidean norm for momentum, not the whole four-vector, but that shouldnโ€™t be a problem โ€“ it makes it easier.โ€™ Then pause, and add, โ€˜In Minkowski space, obviously.โ€™(beginfootnote)Minkowsi space is like Euclidean space but in four dimensions. This might be a good time to add the footnote that there is an even more fundamental equation, which is Einsteinโ€™s field equation of general relativity (Carroll, 2014), but thatโ€™s something to discuss at a later point in time.(endfootnote)

    Also, pure energy = pure nonsense. If Leibniz were somehow able to hear this, he would cackle and turn over in his grave โ€“ if he could muster the energy for it. To be fair, physicists use the word energy all the time, all over the place. Itโ€™s short, sweet, and simple to use on a daily basis, which is fine, just as long as weโ€™re all in agreement about what we mean.

    Energy isnโ€™t fundamental to motion, itโ€™s motion(beginfootnote)Of quantum fields as described by wavefunctions(endfootnote) and interactions giving rise to the construct of energy.

    However, if you do bump into a floating blob of pure energy down the long narrow hall upstairs of your rich auntโ€™s mansion, then, well, yes, that would most certainly be something strange.

    References

    Carroll, S. (2014) Spacetime and Geometry: Pearson New International Edition : an Introduction to General Relativity. 1st. Pearson.

    Einstein, A. (1905a) ‘Zur Elektrodynamik bewegter Kรถrper’, Annalen der Physik, 322(10), pp. 891-921.

    Einstein, A. (1905b) ‘Ist die Trรคgheit eines Kรถrpers von seinem Energieinhalt abhรคngig?’, Annalen der Physik, 323(13), pp. 639-641.

    Eisberg, R. M. and Resnick, R. (1974) Quantum physics of atoms, molecules, solids, nuclei, and particles. New York: Wiley.

    Miller, A. I. (1981) Albert Einstein’s special theory of relativity : emergence (1905) and early interpretation (1905-1911). Reading, Mass ;: Addison-Wesley.

    Okun, L. B. (1989) ‘The Concept of Mass’, Physics Today, 42(6), pp. 31-36.

    Taylor, E. F. and Wheeler, J. A. (1992) Spacetime physics : introduction to special relativity. 2nd ed. edn. New York: W.H. Freeman.


    Featured image: NASA’s Solar Dynamics Observatory captured this image of an X2.0-class solar flare bursting off the lower right side of the sun on Oct. 27, 2014. The image shows a blend of extreme ultraviolet light with wavelengths of 131 and 171 Angstroms. Credit: NASA/SDO. Retrieved 30 Aug 2019, from https://www.nasa.gov/content/goddard/sun-release-x20-class-flare-on-oct-27-2014

    Einstein and Ehrenfest. [Photography]. Encyclopรฆdia Britannica ImageQuest. Retrieved 21 Aug 2019, from https://quest.eb.com/search/132_1510430/1/132_1510430/cite

    Oliver Heaviside (1850-1925) – Science and Society Museum/ Universal Images Group. Oliver Heaviside, English physicist, c 1900.. [Photograph]. Encyclopรฆdia Britannica ImageQuest. Retrieved 1 Sep 2019, from https://quest.eb.com/search/102_541915/1/102_541915/cite

    Hendrik Lorentz (1853-1928) – Science and Society Museum/ Universal Images Group. Hendrik Antoon Lorentz, Dutch physicist, c 1920.. [Photograph]. Encyclopรฆdia Britannica ImageQuest. Retrieved 1 Sep 2019, from https://quest.eb.com/search/102_523268/1/102_523268/cite

    Henri Poincarรฉ (1954-1912) – akg-images / Universal Images Group. Henri Poincare / Photo c. 1890. [Photograph]. Encyclopรฆdia Britannica ImageQuest. Retrieved 3 Sep 2019, from https://quest.eb.com/search/109_171035/1/109_171035/cite

    Joseph J. Thomson (1856-1940) – Science and Society Museum/ Universal Images Group. Sir Joseph J. Thomson, English physicist, late 19th century/early 20th century.. [Photography]. Encyclopรฆdia Britannica ImageQuest. Retrieved 1 Sep 2019, from https://quest.eb.com/search/102_547694/1/102_547694/cite. Cropped by @kjrunia.

    George Frederick Charles Searl FRS(1864-1954) – Royal Society. As printed in Thomson, G. (1955) ‘George Frederick Charles Searle. 1864-1954’, Biographical Memoirs of Fellows of the Royal Society,1, p. 247. Cropped by @kjrunia.

    Wilhelm Wien (1864-1928) – NATIONAL LIBRARY OF CONGRESS / SCIENCE PHOTO LIBRARY / Universal Images Group. Wilhelm Wien, German physicist. [Photography]. Encyclopรฆdia Britannica ImageQuest. Retrieved 1 Sep 2019, from https://quest.eb.com/search/132_1255736/1/132_1255736/cite

    Max Abraham (1875-1922) – Niedersรคchsische Staats- und Universitรคtsbibliothek, Gรถttingen. Max Abraham around 1905. Public domain. Slightly cropped by @kjrunia.

    Albert Einstein, Swiss-German physicist. [Photograph]. Encyclopรฆdia Britannica ImageQuest. Retrieved 22 Aug 2019, from https://quest.eb.com/search/132_1510416/1/132_1510416/cite

    Paul Dirac. [Photography]. Encyclopรฆdia Britannica ImageQuest. Retrieved 24 Aug 2019, from https://quest.eb.com/search/132_1254852/1/132_1254852/cite


  • Einstein’s special relativity in under 6.999 minutes for people on the move

    Einstein’s special relativity in under 6.999 minutes for people on the move


    In 1905, Albert Einstein published an article on moving bodies and electrodynamics. He noticed that Newton’s mechanics of moving bodies weren’t compatible with Maxwell’s equations of electromagnetism. In this article, he reconciled the two by modifying the first. These ideas and mathematical derivations became what we now know as Einstein’s special (theory of) relativity. In this post, we will describe some of its important bits. We start with two fundamental propositions. For the geeks, we will end with answering why special relativity is special.

    Einstein’s postulates

    His first postulate is basically that, whether or not you’re standing on a moving object, such as a ship, a train, or in your car, the same laws of physics apply. If you’re standing on a platform at the train station and you throw a ball in the air, the laws of physics ensure your ball comes down again. If you’re standing in a moving train, that ball still comes down again because the same laws of physics apply. The fact that you’re in motion doesn’t change anything to the rules of the Universe.

    Einstein’s second postulate is basically that, by extension of the first one, the speed of light is the same for everyone. He recognised that Maxwell and colleagues correctly describe light as an electromagnetic disturbance propagating according to the electromagnetic laws of physics. He also recognised that the velocity of the light source plays no role in Maxwell’s equations. And so, if the first postulate is correct, then irrespective of the velocity of an observer relative to the light source, light travels at the same speed $c$ ($c$ is about 300 000 000 m/s) and nothing can go faster.

    An image of the interior of an underground train. Passengers are sitting across each other.
    Whether you’re on the train or standing on the platform, the laws of physics are the same. The propagation of light is a law of electrodynamics not involving the velocity of its source. Hence, its velocity is the same on the train as on the platform, irrespective of its source or the train’s velocity.

    Intuition works mostly, just not really

    Suppose, you’re standing on a train station’s platform. A train passes at a speed of 30 metre per second. From inside the train, our friend throws a tennis ball out the window but in the direction of where the train is headed, at 2 metre per second, right at you. You catch it. At what speed does the ball hit your hand?

    Intuitively, you might say, that’s 30 + 2 = 32 metre per second. This way of calculating is very useful, most of the time. Instinctively, you would simply add the train’s speed and the throwing speed together. This is how Isaac Newton(beginfootnote)And Galileo Galilei before him as this is the so-called Galilean transformation.(endfootnote) would want you to do it. And, mostly, he’s not wrong. Except, well, he is kinda.

    Let’s wonder what would happen if our friend didn’t throw a ball, but, instead, switched on a pocket torch. The train still passes at 30 m/s. Light, however, flies out the torch at a speed of about 300 000 000 m/s. You lift up your hand. It ‘catches’ the light. At what speed does it hit your hand?

    You might say, that’s 30 + 300 000 000 = 300 000 030 m/s. But no. That’s wrong. Remember Einstein’s second postulate? Irrespective of the motion of the observer, light always travels at 300 000 000 m/s and nothing goes faster, full stop. So, by our simple addition, we would have invented a way for light to go faster than light! We would be Nobel Prize winners, surely. Except, it doesn’t, and we’re not.

    Something’s gotta give

    So, if the speed of light is the same to our friend, on the moving train, as it is to us, standing on the platform (do read this again and realise how bonkers this is), then how the Helheim does the Universe achieve this? After he did some relatively simple mathematics โ€“ which a student in secondary education can do โ€“ Einstein realised that something was up with metres and seconds. He proved that what is a metre to us isn’t a metre to our friend and vice versa. Furthermore, what is a second to us isn’t a second to our friend either.

    Basically, since we, on the platform, measure light to be going at 300 000 000 metre per second, and so does our friend on the train, well, that means that our understanding of what metres and seconds are is wrong.

    The point

    Here’s what’s happening. If two ‘things’, people, ‘objects’, or reference frames as physicists call them, move with respect to each other, weird things happen to space and time. Yes, the actual space and time. They are weird. We thought they were just there. Static. Always and everywhere the same. Two unchanging entities. Well, they’re not.

    In the Dutch town of Leiden, the project Leiden Wall Formulas have scattered physics equations throughout the town centre. This is the Lorentz contraction, aptly placed alongside a train track: it describes how space (in this case, a one-dimensional length) is contracted, according to Einstein’s special relativity. Click here for Google Street View.

    Suppose, we are on the train this time. We are in motion relative to our friend on the platform. Our friend then observes that space along our direction of motion becomes smaller, it contracts. They will actually measure our train to be shorter as compared to when it was standing still relative to our friend. They will also observe that our time is being stretched, i.e. our clock slows down. If one second goes by on their clock, they see only 0.9999999995โ€ฆ seconds have past on ours(beginfootnote)This number is a metaphor. The real time difference is too small for my calculator to show.(endfootnote).

    And to us, being on the train, it’s our friend who is travelling (backwards!) relative to us, in fact, the whole world is travelling relative to us, backwards. So, indeed, we measure the world to be shorter in the opposite direction of our motion as well as their time being slowed down.

    Twin paradox

    Now, you might say, hold on: if both of us see each other’s clock slow down, then surely, there is no difference between our clocks. However, when we ride back to our friend, stop at the platform, and compare our clock to our friend’s clock, we do see that our clock is behind. This is the so-called twin paradox. If both can state the same thing, how do their clocks still differ in the end, causing one half of the twin (on the train) to be younger than the half who stayed behind (on the platform)?

    This is due to the switching of reference frames: we were on a moving frame (the train), switched to a moving frame in the opposite direction (returning to our friend), and, lastly, switched to the platform’s frame, standing next to our friend, to compare clocks, while our friend never switched โ€“ he stayed on the platform(beginfootnote)Contrary to many popular and even introductory physics texts, this has less to do with acceleration, even though this does plays a role โ€“ without it, in the real world, one wouldn’t be able to switch reference frames. However, mathematically, acceleration isn’t necessary for solving the so-called twin paradox, switching frames is.(endfootnote). So, the situation isn’t symmetric.

    Length contraction and time dilation. They’re not illusory, they’re real. Many experiments showed that space contracts and time dilates for things in motion as observed by things with a different motion.

    Newton vs Einstein

    To summarise in a slightly more mathematical way โ€“ Newton taught us that one platform metre equals one train metre and that the same is true for seconds:

    1 platform metre = 1 train metre,
    1 platform second = 1 train second.

    Einstein, however, taught us that:

    1 platform metre $\equiv \dfrac{1\text{ train metre}}{\gamma}$,
    1 platform second $\equiv \dfrac{1\text{ train second}}{\gamma}$.

    This $\gamma$ (Greek letter gamma) is crucial. It’s a factor necessary to make sure that the speed of light doesn’t get any faster than the speed of light, even if it’s on-board a moving train. This factor is called the Lorentz factor(beginfootnote)Named after Hendrik Lorentz. The expression for his Lorentz factor is as follows: \[ \gamma = \dfrac{1}{\sqrt{1-\dfrac{v^2}{c^2}}}, \] where $v$ is the speed of the other relative to us and $c$ is the speed of light.(endfootnote).

    We never notice these things though. Usually, our speeds are way to slow for the effects of special relativity to be noticeable. Except for things that do go fast. Without Einstein’s special relativity, GPS satellites wouldn’t work(beginfootnote)Of course, we also need Einstein’s General Relativity for that, but that’s for another time.(endfootnote). Research institutes such as CERN and Fermilab also need to take special relativity in account for the high-energy, fast-flying particles.

    So, our intuition (Newton) is mostly just fine though not precisely right.

    What’s so special about special relativity

    This is a somewhat technical question, requiring a somewhat technical answer, our apologies. Contrary to what is being said in many popular science texts, special relativity isn’t really about constant speeds as opposed to general relativity dealing with acceleration. In fact, special relativity is able to deal with acceleration. It’s just that it works fine as long as we’re assuming things exist in so-called Euclidean space, adhering to Euclidean geometry(beginfootnote)Named after Euclid.(endfootnote). However, after ample deliberation, Einstein concluded through his general relativity that we’re not living in Euclidean-geometric reality at all. We’re living in a Riemannian manifold(beginfootnote)Named after Bernhard Riemann.(endfootnote). In short, Euclidean space is a special case of the more general Riemannian manifold(beginfootnote)If we’re being precise, Hermann Minkowski developed a modified version of Euclidean space which we now call Minkowski space, while Riemannian manifold should actually be named pseudo-Riemannian manifold but many physicists simply call this Riemannian manifold anyway, perhaps because they’re not mathematicians.(endfootnote). Hence, we have special relativity as opposed to general relativity.


    Featured image by StockSnap from Pixabay, modified by @kjrunia (added equations).


  • Rainbows: Alexander’s band

    Rainbows: Alexander’s band


    Two rainbows. The air underneath the primary is brighter than the air in between the primary and the secondary rainbow.
    Figure 1

    If you haven’t yet seen it, next time you see a rainbow, you will look for it: an intense difference between the inside and the outside of the primary rainbow. Also, next time you see a rainbow, you will know why this is. Not always as visible but certainly present if there are enough water drops to go around and the intensity of light is sufficient, there will be a secondary rainbow. The darker space between the primary and secondary is called Alexander’s (dark) band.1


    Just in case you didn’t know, there are three requirements for a rainbow. The Sun should be shining behind you. There should be raindrops in front of you, be it thousands of meters up and away from you or even just a few meters (e.g. a lawn sprinkler). And there should be no clouds or anything else in the way between the Sun, the raindrops, and your eyes.

    Refraction and dispersion

    As you may know, light bundels are refracted by any transparent material they encounter. Twice, actually, at the two surfaces they pass through. This is due to the fact that light, being an electromagnetic disturbance, changes the electric properties of the material, which in turn changes the electromagnetic field inside the material, which in turn changes the direction of the light. In a previous article, we delved deeper into the quantum physics of the matter.

    A computer animation of a light bundle entering a prism. At the two surfaces of the prism, where the light enters and exits, the red part of the light is bent less than the violet part of the light, effectively tearing the light bundle apart into a spectrum of red to violet and everything in between โ€“ the colours of the rainbow.
    Figure 2

    It just so happens that red light gets refracted at a smaller angle than violet light, i.e. red light gets ‘bent’ less. Furthermore, the angle between the incident light ray and red light exiting the water drop is always maximally about 42ยบ, while in the case of violet light this is maximally about 40ยบ. In other words, you won’t see violet light exiting the raindrop at 42ยบ โ€“ it’s all red in that region. Everything in between is yellow, green, and blue โ€“ the rest of the rainbow colours. This is depicted in Figure 3. This is the reason why the ‘white’ light from the Sun โ€“ which is rather a blend of all the colours of the rainbow and not at all white โ€“ gets dispersed in a specific order of different colours.

    Figure 3

    Now, suppose millions of tiny raindrops linger in the air in front of you. Depending on a raindrop’s height relative to you, you are only able to see its outgoing light (having been refracted and reflected inside of it) at a specific angle. Some are at a height just right for you to spot only their red light refractions, while others are at a height offering you a view on their violet light refractions.

    In Figure 4, you can see the relation between the (order of the) colours you can see, the different angles at which different colours exit a raindrop as well as the height of the raindrop. When the raindrop is low enough, all colours are mixed again. Red, yellow, green, violet โ€“ they can all exit the raindrop at an angle smaller than 40ยบ. This is the reason why ‘white’ light is being brought about ‘inside’ the primary rainbow.

    Figure 4

    The raindrops don’t reflect light just once, as shown in Figure 3. Sometimes light gets reflected twice inside, as shown in the first two drops in Figure 5. This is how the inverted order of the colours of the secondary rainbow arises. Sometimes light exits the drop but never reaches your eye because they are too high. Some light is ‘lost’ in a sense. This is why the band between the primary and secondary rainbow is extra dark compared to the rest of the sky.

    Figure 5

    Figure 1, photograph of Alexander’s band by Gnangarra under CC BY-SA 3.0 AU.

    Figure 5, primary and secondary rainbow illustration by CMG Lee under CC BY-SA 4.0, adapted by @kjrunia.


    1. The Greek philosopher Alexander of Aphrodisias was the first to mention this phenomenon in one of his commentaries on Aristoteles’ work on meteorology.[]
  • A radioactive smoking gun

    A radioactive smoking gun


    As we discussed in a previous article, microwave ovens don’t destroy atoms โ€“ their radiation simply isn’t ionising radiation, while it’s exactly the ionising stuff that is bad for your cells, such as gamma rays and cosmic rays. The latter reside at the high-energy end of the electromagnetic spectrum. We also mentioned that microwave oven radiation isn’t the same as the radiation present in Chernobyl. Microwave ovens are not radioactive. But when do we call something radioactive then? What is radioactivity?

    Unstable to the core

    Matter is radioactive when the nuclei of its atoms are unstable enough to decay into different types of nuclei, emitting any of the following ionising radiation in the process:

    • alpha rays, a stream of clumps of two protons and two neutrons;
    • beta rays, a stream of electrons or positrons;
    • gamma rays, the higher-energy form of electromagnetic radiation beyond X-rays;
    • neutrino, a particle with the smallest mass.

    Two well-known examples of radioactive material are plutonium and uranium, mostly associated with nuclear power plants and nuclear weapons. Perhaps less known to be radioactive are, for instance, radon-222, lead-210, polonium-210, and potassium-40. We will explain those numbers in the next section.

    So, again, just to be clear, microwave ovens do not pour any of these particles over your food. Nothing gets ‘nuked’. There’s nothing nuclear going on. Your food, however, might very well be radioactive as it’s likely to contain potassium-40.

    A cartoon of an atom

    Atoms consist of three constituents: electrons, protons, and neutrons. Only the hydrogen atom lacks neutrons, the rest is a composite of all three elements. Figure 1 shows a cartoon of an atom. The vague, yellow band represents the electron cloud. The nucleus has been magnified so that its protons and neutrons become visible.

    Figure 1

    What makes one atom different from the other โ€“ say, calcium from potassium โ€“ is the number of electrons, protons, and neutrons. The latter two are called nucleons. The number of protons is most important here: it determines which element of the famous periodic table of elements (type of atoms) we’re dealing with. Potassium has 19 protons โ€“ this is crucial. If it were to acquire any other number of protons, it stops being potassium. Calcium has 20 protons, for example. The number 40 in ‘potassium-40’ means that its nucleus has 40 nucleons. In other words, it has 40 nucleons – 19 protons = 21 neutrons. Mind you, there are also potassium-39 (19 protons, 20 neutrons) and potassium-41 (19 protons, 22 neutrons).

    All these potassium-versions โ€“ same number of protons, different number of neutrons โ€“ are called isotopes. So, while an element (type of atom) has a very specific number of protons, its number of neutrons may differ. Potassium-40 is an example of an isotope of potassium. This and the other aforementioned potassium-isotopes all occur naturally. However, humans are capable of synthesizing another 22 isotopes, no less.

    Radioactive decay

    The nucleus of potassium-40 is unstable. It decays, as it’s called, mostly into calcium-40, which is a stable atom. During the decay, the potassium atom gains a proton and becomes a calcium atom. In the process, a beta particle, i.e. an electron, and an antineutrino are emitted. This is why the decay is called radioactive: it actively radiates stuff. The electron flies off at great speeds and is potentially ionising, i.e. it is capable of knocking another electron from its atom, thereby potentially destroying molecules such as DNA.

    About 0.01% of the potassium mass in our bodies, acquired through food, is of the potassium-40 variety. About 5000 of these atomic nuclei decay every second in an average human body. So, humans possess a radioactivity of 5000 Bq (becquerel, named in honour of Henri Becquerel, who shared the Nobel Prize with Pierre and Marie Curie for their work in radioactivity).

    Bananas are also radioactive as they also naturally contain potassium-40. Its decay rate lies at 14 Bq, i.e. 14 decaying nuclei per second. It won’t set off a Geiger counter. A truck full of them probably would.

    Ionising radiation and the human body

    To assess the extent to which exposure to ionising radiation is likely to have medical consequences, we use units of ‘sievert’ (the symbol is Sv), a measure for the effective dose of ionising radiation in humans.

    The International Commission on Radiological Protection, an independent, non-governmental organisation providing recommendations concerning ionising radiation, assessed that 1 Sv represents a 5% chance of developing cancer.

    Fortunately, we evolved to absorb safe and small amounts of ionising radiation on a daily basis. The talented and clever Randall Munroe, famous for his xkcd.com cartoons, made a wonderful chart ‘with help from Ellen, a Senior reactor operator at the Reed Research Reactor’. We have blatantly copied their brilliant concept in Figure 2 to concisely give you a feeling for the various amounts of ionising radiation (click to enlarge). However, by all means, do also have a look at their original chart. Note that all these diagrams mainly indicate the orders of magnitude โ€“ not exact values as sieverts may vary a little for various human bodies, and sources and locations mentioned. This should, nonetheless, give you an idea how much radiation professionals in the radiation industry are allowed to be exposed to on an annual basis.

    Figure 2

    Smoking hot

    No, this isn’t going to be pretty. If anything, it’s pretty bad, actually.

    Earlier, we mentioned radioactive elements radon-222, polonium-210, and lead-210. They naturally occur in the soil and air. They are also present in and on tobacco leaves and remain there even after processing. Once inhaled, sticky tar assures that these radioactive elements remain in the small air passageways and lungs indefinitely. Moreover, radon-222 happens to decay into said polonium and lead isotopes. Hence, the latter two build up even more. This is also true for secondhand smoke. Together with toxic substances such as tar, arsenic, nicotine, and cyanide, the ionising radiation emitting decaying nuclei of radon-222, polonium-210, and lead-210 increase chances of developing lung cancer immensely.

    According to Little and colleagues (1965, 1967), polonium-210 accumulates in certain ‘hotspots’ in the lungs. Karagueuzian and colleagues (2012) reported an estimated lung dose of 165 mSv per year. To get a feeling of how much ionising radiation a smoker’s lungs (especially the hotspots) receive annually compared to how much ionising radiation a whole body of a professional radiation worker is allowed to be exposed to, see Figure 3. Mind you, this is just about radioactive decay in the lungs. We haven’t even discussed the other toxic compounds. And so, while someone may worry about Wi-Fi signals, mobile phones, and microwave ovens, which, by physics, they oughtn’t, keep in mind, in this Universe, that same physics tells us their smoking habit is really bad. Like, really.

    Figure 3

    References

    Little, J. B., Radford, E. P., Mccombs, H. L. and Hunt, V. R. (1965) โ€˜Distribution of polonium-210 in pulmonary tissues of cigarette smokersโ€™, The New England journal of medicine, vol. 273, no. 25, p. 1343 [Online]. DOI: 10.1056/NEJM196512162732501 (Accessed 9 August 2019).

    Little, J. B., Radford, E. P. and Holtzman, R. B. (1967) โ€˜Polonium-210 in Bronchial Epithelium of Cigarette Smokersโ€™, Science, vol. 155, no. 3762, pp. 606โ€“607 [Online]. DOI: 10.1126/science.155.3762.606 (Accessed 9 August 2019).

    Karagueuzian, H. S., White, C., Sayre, J. and Norman, A. (2012) โ€˜Cigarette Smoke Radioactivity and Lung Cancer Riskโ€™, Nicotine & Tobacco Research, vol. 14, no. 1, pp. 79โ€“90 [Online]. DOI: 10.1093/ntr/ntr145 (Accessed 9 August 2019).

    Featured photo by Julia Sakelli.


  • A little bit: mirror writing

    A little bit: mirror writing


    Sometimes, we will post a ‘little bit’ on here to highlight the main message of a previous article. You may download the images below as they are published under CC BY-SA 4.0. All published ‘little bits’ will be collected and made available for download on the Little bits section of this website.

    This particular ‘little bit’ is based on Mirror, mirror, what’s up with the mirror writing?

    ----- (Dutch translation follows after this English text.) This is a little card with the title A little bit. On the card is the following text: Why do mirrors only flip left and right of letters and words and everything? Why don't they flip stuff upside down? It's not the mirror. It's you. The mirror just reflects what you're showing it, AFTER YOU ROTATED the piece of paper - or yourself - towards the mirror, flipping left and right in the process. Source: opencurve.info ----- In Dutch: Dit is een kaartje met de titel A little bit. Op het kaartje staat het volgende: Waarom verwisselen spiegels van letters en woorden en alles alleen links en rechts? Waarom nooit ondersteboven? Dat ligt niet aan de spiegel, dat komt door jou. De spiegel reflecteert gewoon wat je het toont, nadat je het papier - of jezelf - NAAR DE SPIEGEL TOE HEBT GEDRAAID. Door die rotatie heb je links en rechts verwisseld. Bron: opencurve.info/nl -----
    ----- (Dutch translation follows after this English text.) This is a little card with the title A little bit. On the card is the following text: Why do mirrors only flip left and right of letters and words and everything? Why don't they flip stuff upside down? It's not the mirror. It's you. The mirror just reflects what you're showing it, AFTER YOU ROTATED the piece of paper - or yourself - towards the mirror, flipping left and right in the process. Source: opencurve.info ----- In Dutch: Dit is een kaartje met de titel A little bit. Op het kaartje staat het volgende: Waarom verwisselen spiegels van letters en woorden en alles alleen links en rechts? Waarom nooit ondersteboven? Dat ligt niet aan de spiegel, dat komt door jou. De spiegel reflecteert gewoon wat je het toont, nadat je het papier - of jezelf - NAAR DE SPIEGEL TOE HEBT GEDRAAID. Door die rotatie heb je links en rechts verwisseld. Bron: opencurve.info/nl -----
  • Just a minute: how do polarized sunglasses work?

    Just a minute: how do polarized sunglasses work?


    According to our best understanding of the observable Universe, it is filled with an omnipresent electromagnetic field. Certain perturbations of that field correspond to what we call (electromagnetic) radiation, the visible part of which we call light or light particles, or photons. These disturbances can have specific but differing frequencies, which, when visible, we may perceive as red, yellow, green, blue or violet. Every photon has a frequency: some value is going up and down over time. It turns out that discovering what it is that is changing over time, is the key to unlocking the answer to how polarized sunglasses work.


    Traditionally, it’s useful to mathematically separate the electromagnetic field into two components: the magnetic and the electric subfield. For a slightly more in-depth discussion of this topic, have a look at Why, exactly, do glass and liquids refract light? Geometrically, these subfields are orientated perpendicularly to one another. Have a look at Figure 1.

    Figure 1. A cartoon of the electromagnetic field. It consists of five parts. Their descriptions are given in the main text.
    Figure 1. A cartoon of the electromagnetic field.

    Part a. A three-dimensional electromagnetic field (EM-field) pervades the observable Universe. Here, it is depicted as a finite block but that’s just a cartoony metaphor. In reality, it has the shape of the Universe, and it’s seemingly infinite, or, to be more precise, it’s everywhere you can possibly look.

    Part b. As said before, it turns out to be very useful to mathematically separate the EM-field into two components: the magnetic and electric subfields, depicted here as two planes orientated perpendicularly.

    Part c. When a photon passes through space, this is where the EM-field is disturbed. It is a local change of electric and magnetic values back-and-forth over time. Very important to note: there is nothing in space going up and down or left or right, it is just a cartoon depicting changing values of the respective subfields. The only thing that is actually spanning through space is the trajectory of the photon, depicted by an orange arrow.

    Part d. For our polarized sunglasses, only the electric subfield is relevant, so, we’ve left out the magnetic arrows, just the electric arrows are shown.

    Part e. Of course, no light beam consists of merely one photon. In reality, a bundle of billions of photons are whizzing through space. The orientation of their EM-components will be at all sorts of angles.

    Inside the glasses

    Figure 2. A cartoon of atoms in the polarized filter of sunglasses. On average, they are lined up in a chain in a way as to allow the electron cloud to mainly move up and down, not left and right. The material is said to be vertically polarized.
    Figure 2. A cartoon of atoms in the polarized filter of sunglasses. On average, they are lined up in a chain in way as to allow the electron cloud to mainly move up and down, not left and right. The material is said to be vertically polarized.

    The atoms of polarized sunglasses are lined up in a chain of atoms in such a way that the most wiggle room they have is in the vertical direction as depicted by Figure 2. Incoming photons transfer their energy, through the EM-field, to the wiggling electron cloud, which starts wiggling even more but only in the vertical direction. The latter will activate the EM-field with a vertically orientated electric subfield component, thereby propagating the vertically polarized parts of the incoming light beam.

    Photons with a horizontal polarization, i.e. with a horizontally orientated electric subfield, wiggle the long chains of the sunglasses’ atoms in the horizontal direction. Their energy gets distributed over billions of atoms, horizontally, and is merely dissipated as heat: too low for light propagation. So, basically, these type of photons disappear and the sunglasses heat up a little bit.

    Lastly, photons with an electric subfield perturbation at an angle in between the horizontal and vertical direction will sometimes pass through, sometimes will dissipate.

    Why are polarized sunglasses vertically polarized?

    When light hits a surface, the outgoing or reflected light mostly consists of photons with the same electric orientation as that of the reflecting surface. So, roads and water mainly reflect horizontally polarized light. When you’re navigating a vehicle, you would definitely want to prevent blindness from reflections off of the road or water.

    Pilots

    Computer screens also emit polarized light. If you hold polarized sunglasses in front of it and turn them, at some point, a computer screen’s light will be blocked.

    A pair of polarized sunglasses are held in front of a computer screen. They are rotated at an angle of 90 degrees. The image of the computer screen disappears as the sunglasses turn black, blocking every light coming off of the screen.

    This is also why pilots don’t wear polarized sunglasses: a slight turn of the head would make it impossible to quickly and reliably read vital information off of their instruments. So, despite what expensive brands would like you to believe, a set of polarized sunglasses called something like ‘aviator sunglasses’ are useless in real-life aviation.

    To check if your sunglasses have genuine polarized filters, tilt them in front of a working computer screen. This way you’ll know which pair to leave at home before flying an aircraft.

    Cinema

    Watching a 3D film in the cinema requires a different type of polarized glasses. So, despite what many people may have told you, it’s not that one glass has been vertically polarized and the other horizontally. True, your left eye needs to receive slightly different images from your right eye but this is achieved in a far more ingenious way.

    As you would want the audience to be able to watch the film despite their (sometimes involuntary) head movements, both the film projector and the 3D glasses are cleverly exploiting circular polarization. Otherwise, the moment you would lovingly tilt your head towards your company’s shoulder, a simplistic left-right combination of horizontal and vertical polarization would render any film star on the big silver screen into a vague and flat character.

  • The Eagle has landed

    The Eagle has landed


    Today, fifty years ago, on the Sunday of 20 July 1969, at 20:17 UTC, aeronautical engineer, test pilot, and astronaut Neil Armstrong reported to the Mission Control’s CAPCOM(beginfootnote)capsule communicator(endfootnote) Charles Duke, and transmitted already historic words.

    NEIL: ‘Houston, Tranquility Base here. The Eagle has landed.’

    CAPCOM: ‘Roger, Twanโ€” Tranquility, we copy you on the ground. You got a bunch of guys about to turn blue. We’re breathing again. Thanks a lot.’

    Of course, ‘Twan'(beginfootnote)’Twan’ is (at least) a Dutch version of the French name ‘Toine’, usually given to males.(endfootnote) was a mistake but who cared because humanity had just achieved the greatest accomplishment thinkable.

    This was and continued to be the stuff that dreams were made of. Many decisions to become a scientist and engineer were due to the hardly over-estimable inspiration the landing on the Moon brought to many of us and the next generations. We can only hope that a sufficient percentage of humanity will be allowed by society to continue to perform scientific research and explorations(beginfootnote)And that scientific data won’t be confused with opinions and/or regarded as part of the larger scheme of deception from a hidden agenda of those who really control the world. ??(endfootnote).

    Here is a clip of the landing, courtesy of NASA: ‘It’s a 16mm film clip showing the final forty seconds of descent. It begins when Charlie Duke calls out sixty seconds of fuel remaining, and with the Little West Crater at the bottom of the window. The time-lapse video runs faster than real-time. Audio runs at normal, real-time speed, however.’


    If you have time (three hours and two minutes), by all means, watch NASA’s restored Apollo 11 Moonwalk.

    Science

    While it was indeed an unprecedented achievement of courage, skill, and applied maths and science, the astronauts also carried out several scientific experiments, yielding some very interesting data on our solar system. Dr. Becky Smethurst, astrophysicist and research fellow at the University of Oxford, highlights the following five things we didn’t know before we went to the Moon, in a video on their YouTube channel:

    • the distance to the Moon,
    • the structure inside the Moon,
    • what the solar wind is made of,
    • what the Moon is made of,
    • how the Moon was formed.

    We highly recommend subscribing to their channel with ever informative and entertaining nuggets of knowledge you both knew and didn’t know you wanted to know about that little agitation called Universe. Happy lunar anniversary! To everyone.

    Screenshot of Becky's video. Click to open their YouTube clip.
    Screenshot of Dr. Becky’s video. Click to open their YouTube clip.

    Featured image: Buzz Aldrin, made by Neil Armstrong. Courtesy of NASA.

  • Is microwave oven radiation unhealthy?

    Is microwave oven radiation unhealthy?


    Some say that the radiation inside a microwave oven is bad for our health. And that itโ€™s bad for our food. Itโ€™s uncertain from where these contentions originate exactly. Even though the introduction of the microwave oven(beginfootnote)They were called โ€˜electronic ovensโ€™.(endfootnote) in our homes took place in the 1960s, among some, they never got rid of their unhealthy reputation entirely. In this article, we will have a look at what its radiation is and how that influences food and vitamins. We will then proceed to answer the question: Is microwave oven radiation unhealthy?

    The word ‘radiation’

    Pripyat, near Chernobyl, Ukraine. When we hear 'radiation', we may associate it with the Chernobyl disaster. That is absolutely not at all what microwave oven radiation is.
    Pripyat, near Chernobyl, Ukraine. When we hear ‘radiation’, we may associate it with the Chernobyl disaster. That is absolutely not at all what microwave oven radiation is.

    In physics, โ€˜radiation’ is the emission or transmission of energy in the form of waves or ‘particles’. Not all radiation is a health hazard to our species as we have evolved to be immune in most cases. Radiation emitted by nuclear reactors is dangerous. However, it may not surprise you that we have evolved to withstand the radiation of tea lights.

    While society generally might not care about what physics says โ€˜radiation’ means, this is what we’re going to be using throughout this article.

    Radiation is not always dangerous. There are more things in everyday life than you might think which are forms of radiation.

    Bananas, also those growing in the wild, naturally possess radioactivity. Our species can handle this. It's infinitely more dangerous for other reasons. Don't litter, folks.
    Bananas, also those growing in the wild, naturally possess radioactivity. Our species can handle the radiation. It’s infinitely more dangerous for other reasons. Don’t litter, folks.

    Some examples of sources of radiation: bananas (which are naturally radioactive), magnets, candle sticks, central heating, club and stage lights, any light source for that matter, including your bathroom light, human bodies, microwave ovens, the Sun, the uranium and plutonium rods of a nuclear plant, and furthermore, anything you can see with your eyes either emits or reflects radiation, right here, right now.

    Just look straight into the eyes of your partner, or friend with merits, lying next to you the next morning: whether you want to or not, they have been literally gushing their radiation all over your body, right here, and are still, right now, the whole time. And not in any spiritual or venereal sense, no no, you have been and are being exposed to actual spurts of electromagnetic radiation discharging from their bodies(beginfootnote)Infrared, mostly. Note that our bodies are also radioactive. We emit ‘particle’ radiation too. On average, about 5000 of our atomic nuclei decay every second (5000 Bq) and emit radioactive radiation.(endfootnote), at energy levels literally more than a hundred thousand times higher than microwave oven radiation.

    In fact, in all the examples above, it’s the same type of radiation as microwave oven radiation, called electromagnetic radiation. The difference is that the examples are more than a hundred thousand times more energetic than microwave oven radiation. Except for a big chunk of the Sun’s radiation, and uranium and plutonium rods. Those entail dangerous forms of ionising radiation, and involve more than just the electromagnetic kind.

    Ionising radiation

    The dangerous form of radiation is called ionising radiation. This is the type of radiation many people think of when they hear the word ‘radiation’. Microwave oven radiation isn’t that.

    If incoming radiation has so much energy that it strips one or more electrons away from their nucleus, we call this ionising radiation. An atom which has lost one or more electrons, we consider to be ionised, and so, we now call it an ion.

    Why is this dangerous? Well, our bodies are made of large strings and knots of intertwined atoms. Our skin, organs, cells, DNAโ€”itโ€™s all made up of trillions of atoms. Those atoms are only able to form these large chains and knots because their electrons keep them together this way.

    Thus, if ionising radiation strips away those electrons from their nucleus, then our molecules, cells, DNAโ€”it all falls apart. And especially damage to our DNA is dangerous as this could develop into cancerous growth. Fortunately, our body has evolved to possess certain superpowers, if you will, enabling it to repair damaged cells and even DNA to an astonishing degree.

    Sadly, there are limits. A sufficient blast of ionising radiation may cause damage beyond our bodies’ repair capabilities and thereby may be the cause for cancer.

    Ionising radiation breaks down atoms, thus molecules, thus organic cells. When our body’s repair mechanism is overwhelmed by the amount of ionisation, this may eventually lead to cancer and organ failure. Microwave oven radiation, however, is not ionising at all. Far from it. It is simply not energetic enough. Not by a stretch.

    Examples of ionising radiation are:

    • subatomic particle radiation: such as protons, neutrons, separate or combined to an atomic nucleus(beginfootnote)Also known as alpha particles.(endfootnote) as well as electrons and positrons(beginfootnote)Also known as beta particles.(endfootnote) flying about, aimed at your general direction;
    • high-energy electromagnetic radiation: cosmic rays, gamma rays(beginfootnote)This is what Dr Bruce Banner was exposed to, turning him into what we call a Hulk. But please, don’t try this yourself. Most likely, you’ll die. At best, you might end up looking like former KGB agent Emil Blonsky or General Thaddeus Ross. If you’re unfamiliar with their tragic fates: Universal Misery.(endfootnote), X-rays(beginfootnote)In hospitals, you will receive much less the amount of X-ray radiation than would be dangerous. Your body is capable of repairing any damage, in this case. It still means that one must be careful, hence, only highly-qualified medical personnel should administer X-ray doses.(endfootnote), the higher-energy UV-light.

    Electromagnetic radiation

    Microwave oven radiation is electromagnetic radiation. What is the latter then? In physics, we have the most successful of theories called quantum electrodynamics (QED), which arose in the 1930s. Do watch these amazing, very accessible videos of the genius and Nobel laureate Richard Feynman, who brought major contributions to QED. It’s the first theory within a larger physical framework called quantum field theory (QFT). In short, without QED, we wouldn’t have had electromagnetism-based technology such as microprocessorsโ€”which means we wouldn’t have had TVs, computers, mobile phones, and internet. To understand the theory is to make a few mental steps. In Why, exactly, do glass and liquids refract light?, we’ve mentioned QFT already. We paraphrase the essence down below.

    Space throughout the entire observable Universe is filled with three-dimensional fields. In fact, fields are a property of space. Space without fields does not exist. With space come fields.

    There are many fields. Two of these fields are the electromagnetic field and the electron field.

    We perceive oscillations at specific frequencies in the electromagnetic field as photons, ‘particles’ of light, if you will, sometimes visible light but most of the time it’s invisible light.

    We perceive oscillations at specific frequencies in the electron field as electrons. If we measure themโ€”interact with them using an electric probe in the lab for instanceโ€”we perceive them as ‘particles’. Usually, we speak about them as ‘particles’, even though they’re not.

    Electrons influence the electromagnetic field. The latter influences electrons in return. Photons are oscillating parts of the electromagnetic field, and so, electrons influence photons, while photons influence electrons. However, electrons are only influenced by photons when the latter have specific energy values, not just any energy value.

    Photons, or, the electromagnetic field disturbances caused by a microwave oven, which we call ‘radiation’, do not have the correct energy value to ionise the atoms in our body.

    Luckily, photons emitted by the person lying next to you don’t have the correct energy value to destroy our atoms either, nor do the regular lights in our home, even though they carry a hundred thousand times more energy than those of a microwave oven.

    All this is perfectly calculable. Because maths and physics.

    A schematic depiction of two fields spanning throughout the entire observable universe. Here, they look like two-dimensional planes hovering over one another with a little bit of empty space between them but in reality they are three-dimensional fields pervading all of space. They are completely intertwined with each other, three-dimensionally. Electrons are specific oscillations in the electron field and are here depicted as darker yellow blobs in the yellow-coloured electron field. Photons, light, or electromagnetic radiation (they are all the same thing) are likewise depicted as darker green blobs in the green-coloured electromagnetic field. The lighter-green blobs represent the influence in the electromagnetic field caused by electrons.
    A schematic depiction of two fields pervading through the entire observable universe. Here, they look like two-dimensional planes hovering over one another with a little bit of empty space between them but in reality they are three-dimensional fields pervading all of space. They are completely intertwined with each other, three-dimensionally. Electrons are specific oscillations in the electron field and are here depicted as darker yellow blobs in the yellow-coloured electron field. Photons, light, or electromagnetic radiation (they are all the same thing) are likewise depicted as darker green blobs in the green-coloured electromagnetic field. The lighter-green blobs represent the influence in the electromagnetic field caused by electrons.

    How do we know?

    Max Planck, a German theoretical physicist (1858-1947), found a way to calculate the energy values for photons. Einstein subsequently used Planck’s formula to come up with another formula allowing us to calculate if atoms would become ionised by certain forms of radiation. In The formula that got Albert Einstein the Nobel Prize and should stop us getting sunburn all the time, we discuss Einstein’s groundbreaking work in quantum mechanics which would later develop to become quantum electrodynamics (QED). Warning: mathematical equations are given in that article.

    This cheery-looking fellow was a physicist, a genius, and a Nobel laureate. He was one of the founders of quantum mechanics. His name was Max Planck. This is a photograph from 1933.
    This cheery-looking fellow was a physicist, a genius, and a Nobel laureate. He was one of the founders of quantum mechanics. His name was Max Planck. This is a photograph from 1933.

    And so, after many more contributions by really clever people, the fascinating branch of science arose through which we are now able to harness the power of electromagnetic radiation, including that of the microwave oven as well as, incidentally, radio signals, TV signals, Wi-Fi, and mobile phone signals.

    If you want to know where in the ‘spectrum of danger’ microwave oven radiation lies, have a look at this diagram. Hint: if radiation in a microwave oven were dangerous, then the lights in your toilet would liquidate you instantly to a warm, sliding, sneakers-covering pulp of mashed guts, lung pudding, and brain leftovers. Also, Planck, Einstein, and others, would be turning over in their graves.

    Heat

    Knowing this, the natural thing to ask is, what about all the heat? If microwave oven radiation is really that low-energy, how does it manage to cook my food boiling hot? The answer is friction.

    Remember that time when you bent a piece of steel wire back and forth quickly for a while? And that the bend eventually became very hot? That’s because you had been moving many molecules back and forth quickly enough to have them heat up the wire due to friction. You don’t need life-threatening amounts of energyโ€”merely the energy of your armsโ€”to make something really, really hot. Microwave radiation does that mainly with the water molecules in food.

    Under the influence of the electromagnetic field inside the microwave, the slightly polarised water molecules rotate back and forth about 2 400 000 000 times per second. Friction with their surroundings cause heat.

    It’s like rubbing your hands together the same number of times per second, which entails friction and causes heat. This is why it’s easier to heat up solid food in the microwave oven than liquids, such as a cup of water: in the latter, water molecules experience less friction than in the first.

    The radiation does nothing to the composition of the atoms. It merely causes molecules to move. Just as a flame does. Or a conventional oven. Making molecules move, that’s all there is to it.

    This also means that you shouldn’t put your hand in an operating microwave oven. Your molecules will start moving too, just like in a conventional oven or when you would put your hands into a flame on the hob. Though, admittedly, in a microwave oven that would mostly only be your water molecules, while in a conventional oven or on the hob, all of your molecules are affected.

    Vitamins

    Does microwave oven radiation destroy vitamins? As stated before, the radiation isn’t ionising, so it does not destroy vitamins. Heat does, however. Just as flames and conventional ovens heat up food and through that heat destroy vitamins, so does radiation. Not because of the radiation but because of simple heat.

    If microwave oven radiation did destroy vitamins, then the light hanging over your dinner table should obliterate the entire dish in an instant. That would be overdoing the concept of having a quick meal, somewhat.

    Here’s a silver lining. The longer food is exposed to heat, the more vitamins are destroyed. So, if there would be a means to heat food as quickly as possible, less vitamins would be destroyed. It might just be that, depending on all kinds of settings and situations, heating food quickly inside an efficient microwave oven spares more vitamins than a slow burn on the stove.

    We see a gas stove. Flames are engulfing the bottom of a pot. Vitamins are destroyed by heat, not microwave oven radiation.
    Vitamins are destroyed by heat, not microwave oven radiation. Beyond a certain temperature, a certain number of vitamins per second are starting to be destroyed. The longer it takes to heat up food after that, the more vitamins are broken down.

    Moreover, cooking food in a pot in its liquids and then throwing away the liquids equals throwing away the dissolved vitamins in those liquids. This happens a lot while cooking on a conventional hob. In a microwave oven, however, everything stays on the same plate. So, even if vitamins have dissolved in the food’s liquid, you’d still have them on your plate.

    Lastly, while microwave ovens don’t emit radiation that is carcinogenic, food itself may very well be, especially when it’s burnt. So, particularly when cooking on the flames and in a conventional oven: don’t burn your food. You know this. Don’t burn it to a crisp and then eat it.

    Incidentally, microwave oven radiation is very unlikely to burn food on a plate(beginfootnote)Which is why many don’t like it since, more often than not, a little bit of that browned-burnt-y flavour does taste good.(endfootnote).

    Is microwave oven radiation unhealthy?

    Courtesy of quantum physics, microwave oven radiation is not unhealthy and by itself, it doesn’t destroy vitamins (heat does). There is also no residual effect that might be dangerous: it’s like switching the light on and off, only a hundred thousand times less energetic than light. One might mix things up with the residual effect of a nuclear bomb or a nuclear disaster. This is due to the billions of heavy atomic nuclei having been hurled into the environment which in turn emit ionising radiation for thousands of years. Microwave ovens don’t hurl atomic nuclei into your food. That would have been deadly, indeed.

    Any text online, in books or magazines, stating otherwise is basically fighting against the quantum mechanical facts of the Universe. And a grumpy Einstein. In which case, other forces and motivations must have been at play in stating these uninformed propositions.

    So, if someone is telling you microwave ovens are unhealthy, ask them for the exact quantum mechanical equations supporting their claim. Sorry but it is what it is, and it is simply this Universe. If quantum mechanics weren’t correct, they’d have never been able to read their uninformed online source, anyway. Computers wouldn’t work. Their mobile phone would be nothing more than a fancy brick. Internet would never have existed. They would never be able to spread false rumours online about microwave ovens if physics were incorrect. Heck, they themselves wouldn’t even exist.

    We see a shirtless man in a canoe. Don't do this without sunscreen. This is dangerous. Not microwave oven radiation.
    Don’t do this without sunscreen. This is dangerous, not microwave oven radiation.

    By the same science, however, do watch out for the Sun’s UV-light this summer. And avoid at all cost, cosmic rays, gamma rays, and particle beams(beginfootnote)Unless exposure takes place briefly, as conducted by awesome and life-saving medical professionals.(endfootnote), so, whatever you do, do not stroll outside your international space station without a protective suit. And don’t eat yellowcake. Ever. Instead, radiate some green beans in a microwave oven. And eat a radioactive banana. Much healthier.


    Photo of Pripyat, near Chernobyl, Ukraine by ะ”ะตะฝะธั ะ ะตะทะฝะธะบ from Pixabay.

  • Why, exactly, do glass and liquids refract light?

    Why, exactly, do glass and liquids refract light?


    Summer has arrived, and you have been served a gorgeous-looking cocktail. Condensation droplets on the glass reveal you are set for a much needed particularly refreshing indulgence. However, just as you were about to soak up the colourful fluid of blissful gratification, your shockingly intelligent child asks why the straw seems to be broken inside your drink. And if not about that, then it’s about why this bear’s head is in the wrong place. Sure, the answer is light refraction, but why, exactly, do glass and liquids refract light?

    People looking at a bear in his habitat in the zoo. The side is transparent, so people can see the bear standing in the water from the side, partially submerged. Due to the light refraction caused by the water and the glass, the bear's head is located at a different place than his submerged body. Dramatically displaced.

    We will provide you with the answer. However, before we begin, we need to ask, TL;DR? Rather not see formulas? Scroll down to the last section, the Quick summary. More curious? Then by all means, read on. I promise, not a single calculation will be done. And if you do read on, you will know actual physics. Shockingly more than most.

    For your convenience, here’s a little table of contents:
    A few incorrect explanations
    Why are they incorrect?
    Step 1. Not particles, not waves: it’s all fields
    Step 2. Maxwell’s field equations
    Step 3. Draw the vectors
    Step 4. The electric field inside of materials
    Quick summary: why, exactly, do glass and liquids refract light?


    A few incorrect explanations

    What would you answer? Here are just a few bad examples which other people (but not you) tend to tell their offspring.

    1. Light takes the fastest route. As its speed differs per material, it needs to change direction. Or: light takes the path of the least amount of action. Same reasoning.
    2. When light enters the glass and the liquid, it bounces back and forth between the molecules and atoms of the material. Due to their crystalline or liquid arrangement, the overall direction of light changes. Hence, light is refracted.
    3. Light consists of particles, so-called photons, which get absorbed by the atoms of the material, causing their electrons to temporarily increase their orbital radius around the nucleus. The instant they fall back to their original orbit, they emit another photon in a direction which depends on and is consistent with the type of atom, i.e. material. The overall result is that the beam of particles has changed direction. Hence, light is refracted.
    4. Huygens’ Principle. Light is a wave. Every point on its wavefront can be a source for a circular wavelet. Draw them, connect the dots and you’ll see: light gets refracted.

    Why are they incorrect?

    1. Okay, this is not incorrect, however, while light does that, it doesn’t explain what really happens. It’s an answer to a different kind of question. So, to be ‘that person’ here, in terms of answer-to-the-question-asked, it’s incorrect after all.
    2. By this logic, light should appear much more spread out due to the probabilistic nature of the supposed bouncing back and forth between chaotically moving or vibrating molecules and atoms. The specific direction of bouncing light is not guaranteed to be as consistent as we nevertheless observe in the real world. The resulting image should be a blur. It is not.
    3. Here too, light should appear much more spread out. The direction of the re-released photon is not guaranteed to be in the direction we observe in the real word. A photon could be re-emitted in any direction, regardless of the type of atom. The frequency of the photon correlates with the atomic configuration, not its direction. Moreover, ‘getting absorbed’ and ‘re-emitted’ are not well defined. What does that even mean?
    4. This is a sophisticated one. At first glance, it does produce an angle for the outbound light beam. However, Huygens’ Principle only corresponds to observations if you cherry pick from multiple possibilities. See Figure 1 for a brief explanation.
    (a) The vertical lines represent the crests of the light wave. The blue area is the glass or liquid. As light only bends in these materials at an angle, the diagram shows a beam of light approaching the surface of the material at an angle. Huygens proposed that at every instance 'wavelets' (drawn here as segments of dotted circles) can be thought emanating at every point in space, growing over time. Connecting the wavefronts of those wavelets predicts the course of the next wave (crest). As the bottom of the incoming crests hit the surface first, those wavelets will have had time to grow larger before the top of the incoming crests hit the surface. (b) Over time, multiple wavelets can be thought to have developed. (c) Where the wavelets intersect each other wave crests can be drawn. The result seems to be the predicted new progression of the light beam inside of the material. (d) However, over time, multiple intersections will have developed. By Huygens' logic, multiple wave crests could be drawn. This, however, would result in a diffuse light wave, spreading out its light instead of a distinct bending of the one beam. Stating that a situation as sketched in (c) will occur is selectively choosing a preferred scenario while (d) shows multiple would occur. Hence, Huygens' principle seems right at first but ultimately breaks down over time.
    (a) The vertical lines represent the crests of the light wave. The blue area is the glass or liquid. As light only bends in these materials at an angle, the diagram shows a beam of light approaching the surface of the material at an angle. Huygens proposed that at every instance ‘wavelets’ (drawn here as segments of dotted circles) can be thought emanating from the wavefront at every point in space, growing over time. Connecting the wavefronts of those wavelets predicts the course of the next wave (crest). As the bottom of the incoming crests hit the surface first, those wavelets will have had time to grow larger before the top of the incoming crests hit the surface. (b) Over time, multiple wavelets can be thought to have developed. (c) Where the wavelets intersect each other wave crests can be drawn. The result seems to be the predicted new progression of the light beam inside of the material. (d) However, over time, multiple intersections will have developed. By Huygens’ logic, multiple wave crests could be drawn. This, however, would result in a diffuse light wave, spreading out its light instead of a distinct bending of the one beam. Stating that a situation as sketched in (c) will occur is selectively choosing a preferred scenario while (d) shows multiple would occur. Hence, Huygens’ principle seems right at first but ultimately breaks down over time.

    Step 1. Not particles, not waves: it’s all fields

    So, what does make light refract then? We need to take a few mental steps. Here is the first one, which you’ll just have to get used to.

    Space throughout the entire observable Universe is filled with fields. In fact, fields are a property of space. Space without fields does not exist. With space come fields. Points in most fields not only have a value, they also have a direction. They are called vector fields. Some are called scalar fields; their points have no direction, they only have values. There are more types of fields, such as tensor fields and fermionic fields. This is quantum field theory (QFT), the most successful and accurate theory to date. Has been for well over ninety years (including a renaissance in the 1970s).

    Next question is, what concrete fields are we talking about? You probably heard of or read about the Higgs field(beginfootnote)It just so happens this is not a vector field; its points have no direction, just values, and so, it is a scalar field.(endfootnote). In 2012, the Large Hadron Collider at CERN produced an oscillation in the Higgs field or rather an excitation. That excitation is what we call the Higgs particle. The energy produced inside the LHC was more than enough to cause an excitation of the Higgs field, which we perceive as a particle(beginfootnote)The Higgs particle itself was indirectly observed as its lifespan is too short. It decays quickly into other particles, or excitations, in other fields. Those, however, live long enough for the detectors to observe.(endfootnote). The field was proven to be a real thing. Two Nobel Prizes were awarded to Franรงois Englert and Peter Higgs for having proposed the existence of the Higgs field forty-eight years earlier. It proved how humans with their shockingly tiny brains were able to probe the depths of the subatomic world, a thousand times smaller than the atomic nucleus, and the entire observable Universe at the same time. By using maths and, forty-eight years later, by building ingenious experiments.

    There are more fields. There is an electron field. Most of the time, the field has value zero. But when the values of a tiny part of that field oscillate at a distinct frequency, we call that an electron.

    There is also an electromagnetic field. A stream of billions of local oscillations of a range of frequencies is what we call a beam of visible light. It’s practical to sometimes talk about it as it being particles (called photons) as well as it being waves (electromagnetic radiation). It depends on what you’re calculating.

    You could say there is also a proton field, although there are more fundamental fields than this, for instance the quark and gluon fields (protons aren’t elementary particles, they consist of quarks and gluons). However, for the purpose of this post, we will work with the simpler notion of a proton field. A fairly local oscillation is a proton, which we usually perceive as a particle.

    There are many more fields but to discuss them all would justify a separate article. Or several books. And a couple of years of study.

    So, what is light, what are electrons, what are protons or quarks? Are they both particle and wave? No, that’s an old and misleading question. Are they sometimes particles, sometimes waves then? No, also not that.

    ‘Particles’ aren’t actual particles like tiny silver ball bearings or something like that. They are best described as a mathematical function, which we call the wave function (denoted by the symbol ฮจ). When measured they are fairly local oscillations or excitations at specific frequencies in fields pervading through all of space, almost behaving like particles. Sometimes, it’s practical to mathematically model them as either particles or waves, depending on the situation. However, it’s meaningless to state they are either or both at the same time. It’s more accurate to just treat them as mathematical wave functions instead of anything elseA personal conviction is currently that the wave function is all there is. Elementary ‘particles’ are wave functions. Nothing more, nothing less. Favouring ‘tangible objects’ over ‘mere mathematical descriptions’, which, nevertheless have been proven to be incredibly accurate after billions and billions of experimental runs, is really just exposing our limited understand of quantum physics as confined by everyday, large-scale experiences such as playing with base-, basket- and footballs. There is no reason, however, to assume the latter are a measure to gauge the subatomic foundation of our Universe. In my view, that should be the wave function. โ€” KJ.

    Figure 2. Three fields of space are drawn stacked. In reality, they are three-dimensional and not stacked and separated as depicted here. Instead, they are occupying the same space, completely blended with each other.
    Figure 2. Three fields of space are drawn stacked. They are two-dimensional here, but in reality they are three-dimensional and fill the same three-dimensional space, completely immersed in and blended with each other. The problem is that drawing mixed and blended three-dimensional stuff is hard on a two-dimensional screen. In this diagram, no ‘particles’ are present at the moment. There are no oscillating excitations in the fields. In other words, the field values are zero, there are no particles, but the fields are still there. Filling space. (Just to be entirely precise, in reality, the fields do always oscillate a little bit as predicted by Heisenberg’s uncertainty principle.)

    Step 2. Maxwell’s field equations

    James Clerk Maxwell was, besides Scottish, a scientist in the field of mathematical physics. Having studied the previous work of Faraday, Gauss, and Ampรจre, he showed that an electric field and a magnetic field were the same thing, just different aspects of it. That thing is what we now call the aforementioned, space-filling electromagnetic field. He also showed that light was an electromagnetic phenomenon. A disturbance in the field.

    Engraving of James Clerk Maxwell by G. J. Stodart from a photograph by Fergus of Greenock. Frontpiece in James Maxwell, The Scientific Papers of James Clerk Maxwell. Ed: W. D. Niven. New York: Dover, 1890. Public domain.

    He formulated a set of four differential equations. This set bears his name. We will only ‘use’ two of the four:

    \begin{align}
    \mathbf{\nabla} \cdot \mathbf{E} &= \frac{\rho}{\varepsilon_0}, \\
    \mathbf{\nabla} \times \mathbf{E} &= -\frac{\partial \mathbf{B}}{\partial t}.
    \end{align}

    I say, ‘use’, but don’t worry, we’re not going to do any complicated calculations.

    Equation (1) is called Gauss’s law and shows how the electric field (which is one aspect of the electromagnetic field), denoted by E, is influenced by a charge $\rho$, such as the negative charge of an electron or the positive charge of a proton. The symbol $\varepsilon_0$ denotes a constant, which differs depending on the material. The subscript 0 denotes it’s the constant of the vacuum of space. This physical constant $\varepsilon_0$ has different names such as vacuum permittivity, permittivity of free space or the electric constant. In this article, we will use different values for this constant, however. We will use

    \[ \varepsilon_\text{air} \text{ and } \varepsilon_\text{mat}. \]

    ‘Mat’ is short for ‘material’ which could be glass or liquid, for example. The precise numerical values we won’t use, because that’s not important for understanding why light bends. These two epsilons will turn out to play a pivotal role in the bending of light by materials, however. Do read on, I’d say.

    In Figure 3, the three fields in space are again depicted. This time, you can see the elevated values as blobs in the proton field. They are protons. They are surrounded by electron blobs as is depicted in the electron field. Both ‘particles’ influence the electromagnetic field, or, rather, the electric subfield thereof. Note that the blobs in the latter do not constitute ‘particles’, merely influences in the electric field.

    Figure 3. Protons and electrons, together constituting atoms, influence the electromagnetic field.
    Figure 3. Protons and electrons, together constituting atoms, influence the electromagnetic field.

    Step 3. Draw the vectors

    Have a look at Figure 4. The orange arrow or vector denotes the direction of the light inside whatever material we have, such as glass or a liquid. Maxwell showed that light, being an oscillation in the electromagnetic field, has an oscillatory component in the electric subfield of the electromagnetic field. And that electric field is orientated perpendicular to the direction of light. This is represented by the green vector.

    Figure 4. Light inside of the material falls at an angle onto the surface of the material. It has an electric field oscillation perpendicular to its direction.
    Figure 4. Light inside of the material falls at an angle onto the surface of the material. It has an electric field oscillation perpendicular to its direction.

    It is important to note that the electric field vector has two fundamental components, namely a component vector parallel to the surface, denoted by the symbol $\parallel$, and a component vector perpendicular to the surface, denoted by the symbol $\perp$. This is depicted in Figure 5.

    Figure 5. The electric field inside the material, caused by the light, has two vector components: parallel and perpendicular to the surface.
    Figure 5. The electric field inside the material, caused by the light, has two vector components: parallel and perpendicular to the surface.

    Exactly at the surface, the transition from the material to air, the electric field of the material and the electric field of the air will have to ‘slide’ to an equal value (or else we would have a tear in our universe). This means that

    \begin{align}
    \varepsilon_\text{air}(\mathbf{\nabla} \cdot \mathbf{E}_\text{air}) &= \varepsilon_\text{mat}(\mathbf{\nabla} \cdot \mathbf{E}_\text{mat}), \\
    \mathbf{\nabla} \times \mathbf{E}_\text{air} &= \mathbf{\nabla} \times \mathbf{E}_\text{mat}.
    \end{align}

    If we do a little bit of calculus, we come to the following equations:

    \begin{align}
    \mathbf{E}_{\text{mat}\parallel} &= \mathbf{E}_{\text{air}\parallel} \\
    \varepsilon_\text{mat}\mathbf{E}_{\text{mat}\perp} &= \varepsilon_\text{air}\mathbf{E}_{\text{air}\perp}.
    \end{align}

    So, the vector components of both air and the material parallel to the surface are equal. However, the vector components perpendicular to the surface are not. This is the crux:

    \[ \varepsilon_\text{mat} \neq \varepsilon_\text{air}. \]

    In fact,

    \[ \varepsilon_\text{mat} > \varepsilon_\text{air}. \]

    This means that the perpendicular vector component of air has to be larger than that of the material in order to satisfy equation (6). This is depicted in Figure 6.

    Figure 6. Because the epsilon (electric constant) of the material is larger than that of air, the perpendicular component vector of air has to be larger due to satisfy the equation. Here, a new resultant vector of the electric field in air has been drawn superimposed on the old electric field vector to emphasise the difference.
    Figure 6. Because the epsilon (electric constant) of the material is larger than that of air, the perpendicular component vector of air has to be larger due to satisfy the equation. Here, a new resultant vector of the electric field in air has been drawn superimposed on the old electric field vector to emphasise the difference.

    The only thing left to do, is to draw the new direction of the light in the air. As we know that the direction of the electric field is perpendicular to the direction of light, courtesy to Maxwell and colleagues, we can easily construct the new course of the light outside in the open air as is depicted in Figure 7.

    Figure 7. The new direction of light in air is refracted relative to the original angle inside of the material, just as vector calculus predicted.
    Figure 7. The new direction of light in air is refracted relative to the original angle inside of the material as predicted by vector calculus, just as observed in reality.

    Step 4. The electric field inside of materials

    The question is now: what causes the electric constant of materials to be so different?

    Figure 8 shows a schematic depiction of what light, being the cause for disturbances in the electric field itself, does to electrically charged ‘particles’ inside a material in terms of quantum field perturbations. Note how the alignment of charges and thus the oscillations in the electromagnetic field have changed in such a way that the electric subfield as a whole, inside of the material, has to have changed values as well. These changes are encapsulated in the electric constant, $\varepsilon_\text{mat}$.

    Light changes a material’s electromagnetic configuration, which then influences the trajectory of that same light.

    Figure 8. Light changes the electromagnetic configuration inside a material, which then influences the trajectory of that same light.
    Figure 8. Light changes the electromagnetic configuration inside a material, which then influences the trajectory of that same light.

    Quick summary: why, exactly, do glass and liquids refract light?

    The Universe, i.e. space itself(beginfootnote)With ‘space’, we don’t mean ‘outer space’ but rather the thing we move in, the volume, the expanse, the invisible yet essential thing allowing us to move back-and-forth, up-and-down, left-and-right.(endfootnote), contains an omnipresent electromagnetic field. Mathematically, we can divide this field up into two components: its electric (sub)field and its magnetic (sub)field.

    Electrons in glass and liquids as well as light are influenced by the electric field. At the same time, they influence that same electric field. When light hits the material, it changes the electric field inside the material. This makes electrons bring about opposite electric field changes in turn. The net electric field inside the glass changes the light’s direction of propagation in a perfectly predictable way. Courtesy of vector calculus.

    Or, if you prefer:

    Light pushes on electrons via the electric field. Electrons push back a bit via the same field. Light says, ‘Okay, okay, relax!’, and takes a slightly different route.

  • Finding the normal force in planar non-uniform circular motion using polar coordinates

    Finding the normal force in planar non-uniform circular motion using polar coordinates


    In this post, we will derive an expression for the normal force on a uniform mass which is in planar non-uniform circular motion using polar coordinates. Finding this expression is enormously useful to calculate under which circumstances a mass would be slung off its orbital path. Of course, there are numerous situations for which we should be able find the normal force. Here, we will look at a system as shown in Figure 1. Sometimes, obtaining an expression in terms of the variables given is not straightforward. You will find a useful trick in step 7 to arrive at an expression in terms of a simple $\theta$ instead of its secondary-order derivative $\ddot\theta$ which we initially obtain.

    This could be seen as an undergraduate-physics-level post. Download this article


    Notation

    We will apply Newton’s notation (the dot notation) whenever possible as this is the most compact form. For instance, if $\mathbf{x}$ is a vector, then its first-order and its second-order derivative with respect to time $t$ are denoted by

    \[ \dot{\mathbf{x}}\text{ and }\ddot{\mathbf{x}}, \]

    respectively. Where needed, in order to state explicitly that we are dealing with a time-derivative and to help in solving a time-integral for example, we will use Leibniz’s notation, i.e.

    \[ \frac{\text{d}\mathbf{x}}{\text{d}t}\text{ and }\frac{\text{d}^2\mathbf{x}}{\text{d}t^2}. \]

    Assignment

    Look at the system as sketched in Figure 1. Imagine we stand in front of this system. Mass $m$ is attached to a model string. At $t=0$, it rests at level with the centre of the cylinder with radius $R$ with the string draped over the top. A constant force $\mathbf{P}$ pulls the string downwards. At a later time $t$, mass $m$ has slid over the top with a coefficient of friction $\mu$. Let $\theta$ denote the angle between its initial and its current position, subtended at the centre of the cylinder. Calculate the normal force on $m$, and, hence, proof that the radius of the cylinder is irrelevant.

    Figure 1. The system

    Step 1. Force diagrams and unit vectors

    It is essential to draw force diagrams and unit vectors to define the acting forces and parameters. We choose the unit vectors to be the radial and the tangential vectors. This makes calculating most forces a lot easier. This is done in Figure 2.

    Figure 2. Force diagram and unit vectors at time $t>0$

    We identify the following forces on $m$:

    • $\mathbf{P}$ is the vector denoting the constant force pulling the model string,
    • $\mathbf{N}$ is the vector denoting the normal force acted on $m$ by the cylinder,
    • $\mathbf{F}$ is the vector denoting the frictional force,
    • $\mathbf{W}$ is the vector denoting the weight of $m$ as a result of the gravitational field of whatever planet the system is located,
    • $\mathbf{e}_r$ is the radial unit vector,
    • $\mathbf{e}_\theta$ is the tangential unit vector.

    Step 2. Apply Newton’s second law

    As this is a dynamical system, where $m$ is in non-uniform circular motion, we apply Newton’s second law, more specifically in the following form:

    \begin{equation}
    \sum\mathbf{F} = m\ddot{\mathbf{r}},
    \end{equation}

    where $\ddot{\mathbf{r}}$ is the rate of change of the rate of change over time, that is, the second time-derivative of the displacement vector $\mathbf{r}$ of mass $m$. We can now easily identify the constituents of the vector sum as we did that already in Step 1. And so, equation (1) becomes

    \begin{equation}
    m\ddot{\mathbf{r}} = \mathbf{P} + \mathbf{N} + \mathbf{F} + \mathbf{W}.
    \end{equation}

    Step 3. Rewrite the forces in terms of their magnitudes and unit vectors

    As pulling force $\mathbf{P}$ with magnitude $|\mathbf{P}|$ acts in the direction of tangential unit vector $\mathbf{e}_\theta$, we can write for $\mathbf{P}$:

    \begin{equation}
    \mathbf{P} = |\mathbf{P}|\mathbf{e}_\theta.
    \end{equation}

    Since we don’t have any other information regarding this force, we leave it at that.

    Normal force $\mathbf{N}$ points in the direction of radial unit vector $\mathbf{e}_r$, so, we write:

    \begin{equation}
    \mathbf{N} = |\mathbf{N}|\mathbf{e}_r.
    \end{equation}

    Friction $\mathbf{F}$ is in the opposite direction of the tangential unit vector $\mathbf{e}_\theta$, so, we need to place a minus-sign in its expression. Furthermore, as (dry) friction is usually modelled by the product of the coefficient of friction and the magnitude of the normal force, we can write:

    \begin{equation}
    \mathbf{F} = \mu|\mathbf{N}|(-\mathbf{e}_\theta).
    \end{equation}

    Lastly, weight is the force due to gravity, $|\mathbf{W}|=mg$, where $g$ is the gravitational constant. However, we need to express this force in terms of its components. In this case, those components are directed parallel to the radial and tangential unit vectors. As the latter are pointed (partly) upwards, as opposed to the downwards-pointing weight, we already know that both its components carry a minus-sign, i.e. $(-\mathbf{e}_r)$ and $(-\mathbf{e}_\theta)$. What remains, is the correct expression for the magnitude of the weight in terms of its respective unit vectors.

    To clearly show how we get an expression for $\mathbf{W}$ in terms of its components along the directions of $\mathbf{e}_r$ and $\mathbf{e}_\theta$, have a look at Figure 3.

    Figure 3. Finding the components of $\mathbf{W}$

    What you see is just the weight vector $\mathbf{W}$ from our force diagram in Figure 2, including the radial and tangential unit vectors $\mathbf{e}_r$ and $\mathbf{e}_\theta$. For visual clarity, we subtended them on mass $m$. Also added are the two component vectors in the opposite direction of the unit vectors for which we need to find expressions.

    Let component vector $\mathbf{v}_r = a(-\mathbf{e}_r)$ and component vector $\mathbf{v}_\theta = b(-\mathbf{e}_\theta)$, where $a$ and $b$ are some magnitude value such that the vector sum of $\mathbf{v}_r$ and $\mathbf{v}_\theta$ equals $\mathbf{W}$. In other words,

    \begin{equation}
    \mathbf{W} = \mathbf{v}_r + \mathbf{v}_\theta = a(-\mathbf{e}_r) + b(-\mathbf{e}_\theta).
    \end{equation}

    To find the values of the magnitude of $a$ and $b$, we use the fact that the magnitude $|\mathbf{W}| = mg$. So, using high school trigonometry, we deduce that

    \begin{align}
    a &= mg\sin\theta, \\
    b &= mg\cos\theta.
    \end{align}

    Now, we can write $\mathbf{W}$ in terms of its components by substituting equations (7) and (8) into (6):

    \begin{equation}
    \mathbf{W} = mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta).
    \end{equation}

    And so, if we substitute equations (3), (4), (5), and (9) into equation (2), we get:

    \begin{align}
    m\ddot{\mathbf{r}} &= |\mathbf{P}|\mathbf{e}_\theta + |\mathbf{N}|\mathbf{e}_r + \mu|\mathbf{N}|(-\mathbf{e}_\theta)\nonumber \\
    &\hspace{2em}+ mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta).
    \end{align}

    Step 4. Express the Cartesian $\ddot{\mathbf{r}}$ in polar coordinates

    As we know that the expression for the second time derivative of non-uniform circular motion is

    \begin{equation}
    \ddot{\mathbf{r}} = -R\dot{\theta}^2\mathbf{e}_r + R\ddot{\theta}\mathbf{e}_\theta,
    \end{equation}

    where $R$ is the radius of the circular motion, i.e. the cylinder. We proceed to substitute this into equation (10).

    And so, we get

    \begin{align*}
    m(-R\dot{\theta}^2\mathbf{e}_r + R\ddot{\theta}\mathbf{e}_\theta) &= |\mathbf{P}|\mathbf{e}_\theta + |\mathbf{N}|\mathbf{e}_r + \mu|\mathbf{N}|(-\mathbf{e}_\theta) \\
    &\hspace{2em}+ mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta),
    \end{align*}

    which, of course, after expansion, becomes

    \begin{align}
    -mR\dot{\theta}^2\mathbf{e}_r + mR\ddot{\theta}\mathbf{e}_\theta &= |\mathbf{P}|\mathbf{e}_\theta + |\mathbf{N}|\mathbf{e}_r + \mu|\mathbf{N}|(-\mathbf{e}_\theta) \nonumber \\
    &\hspace{2em}+ mg\sin\theta(-\mathbf{e}_r) + mg\cos\theta(-\mathbf{e}_\theta).
    \end{align}

    Step 5. Resolve radially and tangentially

    We can now resolve equation (12) into its radial and tangential components.

    \begin{align}
    \mathbf{e}_r &: -mR\dot{\theta}^2 = N – mg\sin\theta, \\
    \mathbf{e}_\theta &: mR\ddot{\theta} = P – \mu N – mg\cos\theta.
    \end{align}

    Step 6. Write down the equation of motion (in polar coordinates)

    Rearranging equation (14), we can write down the second-order differential equation of motion:

    \begin{equation}
    \ddot{\theta} = \frac{P – \mu N – mg\cos\theta}{mR}.
    \end{equation}

    While we could have solved equation (14) for $N$, this would still leave us with the second time-derivative of $\theta$. Instead, we want an expression of $N$ in terms of a simple $\theta$. This means that we need to get rid of $\ddot{\theta}$ in some way. It is not immediately clear how equation (14) or (15) should be operated on to achieve this. However, here is a neat trick.

    Step 7. The trick

    Have a look at the following equation where we apply the chain rule:

    \begin{equation}
    \frac{\text{d}\dot{\theta}^2}{\text{d}t} = \frac{\text{d}\dot{\theta}^2}{\text{d}\dot{\theta}}\frac{\text{d}\dot{\theta}}{\text{d}t} = 2\dot{\theta}\frac{\text{d}\dot{\theta}}{\text{d}t} = 2\dot{\theta}\ddot{\theta}.
    \end{equation}

    So, if we substitute equation (15) into (16), we get

    \begin{equation}
    \frac{\text{d}\dot{\theta}^2}{\text{d}t} = 2\dot{\theta}\left(\frac{P – \mu N – mg\cos\theta}{mR}\right).
    \end{equation}

    If we now integrate both sides with respect to time, we get

    \begin{align}
    \int \frac{\text{d}\dot{\theta}^2}{\text{d}t}\text{d}t &= \int 2\dot{\theta}\left(\frac{P – \mu N – mg\cos\theta}{mR}\right)\text{d}t, \nonumber \\
    \dot{\theta}^2 + A &= 2 \int \frac{\text{d}\theta}{\text{d}t}\left(\frac{P – \mu N – mg\cos\theta}{mR}\right)\text{d}t, \nonumber \\
    &\text{where $A$ is an arbitrary constant}, \nonumber \\
    \dot{\theta}^2 + A &= 2 \int \left(\frac{P – \mu N – mg\cos\theta}{mR}\right)\text{d}\theta, \nonumber \\
    \dot{\theta}^2 + A &= \frac{2}{mR} \int (P – \mu N – mg\cos\theta)\,\text{d}\theta, \nonumber \\
    \dot{\theta}^2 + A &= \frac{2}{mR} \left( P\int 1\,\text{d}\theta – \mu N\int 1\,\text{d}\theta – mg\int \cos\theta\,\text{d}\theta\right), \nonumber \\
    \dot{\theta}^2 + A &= \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2mg\sin\theta}{mR} + B, \nonumber \\
    &\text{where $B$ is an arbitrary constant}, \nonumber \\
    \dot{\theta}^2 &= \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R} + B – A, \nonumber \\
    \dot{\theta}^2 &= \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R} + C, \\
    &\text{where $C=B-A$} \nonumber.
    \end{align}

    Solving the initial condition problem to find $C$, we use the fact that at $t=0$, angle $\theta = 0$, thus $\dot{\theta} = \ddot{\theta} = 0$. This renders $C = 0$ in equation (18), and so, we have

    \begin{equation}
    \dot{\theta}^2 = \frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R}.
    \end{equation}

    Note, we now have obtained an expression for $\dot{\theta}^2$ which already appeared in equation (13). We can, therefore, substitute equation (19) in (13), and we obtain:

    \begin{equation}
    -mR\left(\frac{2P\theta}{mR} – \frac{2\mu N\theta}{mR} – \frac{2g\sin\theta}{R}\right) = N – mg\sin\theta.
    \end{equation}

    Expanding and rearranging this, we get

    \begin{align}
    N – mg\sin\theta &= -2P\theta + 2\mu N\theta + 2mg\sin\theta, \nonumber \\
    N – 2\mu N\theta &= -2P\theta + 2mg\sin\theta + mg\sin\theta, \nonumber \\ N(1 – 2\mu \theta) &= -2P\theta + 3mg\sin\theta, \nonumber \\
    N &= \frac{3mg\sin\theta – 2P\theta}{1-2\mu\theta}.
    \end{align}

    So, now we have an expression of $N$ in terms of the gravitational constant $g$, the variables $m$, $\mu$, and $P$, and the more reasonable $\theta$ instead of $\dot\theta^2$.

    And so, if we want to calculate when a mass would be slung out of its orbital path, we write $N = 0$ as this means, in physical terms, that the mass isn’t resting on the cylinder anymore (since it doesn’t exert a normal force on the mass). In other words, find the roots of equation (21) to find the one unknown variable. Note, $R$ does not play a role. Of course, bear in mind that $m$ is a point mass.

  • Why do wet clothes dry?

    Why do wet clothes dry?


    In the Northern Hemisphere, summer has arrived. The time has come for us to chase the general public with water guns, jump through the neighbour’s garden sprinklers’ water rays, and either carefully place a soggy, wet sea cucumber on a human’s belly during their beach nap(beginfootnote)The author does not approve of this. Sea cucumbers should be left alone.(endfootnote), or simply dump them (the human) in the actually-still-too-cold seawater, especially if you love them. At the end of the day, after all those wet adventures, nothing will beat hanging your clothes out to dry in a soothing breeze of fresh alpine air.

    A few years ago, a friend asked what exactly causes wet clothes to dry. How does that work, exactly? I thought it was a great question because what may seem like a simple problem actually exposes one of the fundamental aspects of the way our universe works and, at the same time, forms one of the main causes for headaches among undergraduates: the second law of thermodynamics.

    TL;DR? Don’t like mathematics (high school level) and prefer to read the ‘dashboard’ version? Skip to the bottom. Everyone else, please read on!


    A box of gas

    Let’s first paint ourselves a simpler picture than the actual situation where we wear whatever is the latest summer catwalk beach fashion. For now, we will also ignore the Sun, we will ignore the wind, and we will ignore the humidity of the air.

    Imagine, your colourful pair of swimming trunks is actually a simple box with a hundred gas molecules. The particles bounce chaotically back and forth against each other and the walls of the box itself.

    Now, let there be a hole in the wall. Imagine, by pure chance, one molecule escaping the box through the hole, arriving in another container of exactly the same size. This obviously means there are now only 99 gas molecules left in the original box.

    Figure 1. Two boxes with gas molecules bouncing around inside. In box (a), one has escaped, 99 remain. In box (b), 98 remain, two have escaped.
    Figure 1. Two pairs of boxes with gas molecules bouncing around inside. In (a), one molecule has escaped to the right box through the hole, 99 remain in the left box. In (b), two have escaped to the right box, 98 remain in the left box.

    Have a look at Figure 1a. Assuming all gas molecules look exactly alike, how many ways do we have to arrange them in order to get the same result? Well, instead of this particular molecule having escaped, any other one of these hundred molecules could have escaped just as well. And so, as each one of the hundred gas particles was capable of escaping the box, exactly a hundred possibilities could have led to the same outcome (i.e. 1 escaped, 99 remain). In other words, exactly one hundred different configurations, or microstates, will entail the microstate of the box where it lost one molecule while 99 remain inside. Let’s call this number $W$. And let’s call that number for the microstate where one molecule escaped (and 99 remain inside), $W(1)$. So, $W(1)=100$.

    Now, imagine not one, but two particles flew out, as is depicted in Figure 1b. Well, this means that a different number of arrangements would have led to this situation or microstate. As concluded above, for the first particle, one hundred possibilities existed as there were as many particles in the box, originally. For the second particle, however, only 99 possibilities existed since one had left the building already! Since for every 100 possibilities for the first molecule, 99 other possibilities exist for the second molecule, we calculate that the total number of possibilities leading to this particular state (i.e. 2 escaped, 98 remain), is $100 \times 99 = 9900$. However, since it doesn’t matter which of the two particles leaves the box first and which second, as they look exactly alike, we can divide that number by two, giving a total number of $4950$ possibilities. And so, $W(2) = 4950$.

    More accurately, in general, to calculate the possible combinations in a situation like this, we use the formula

    \begin{equation} W(k) = \frac{n!}{k!(n – k)!}, \end{equation}

    where $n$ is the number of molecules in the box initially, which is 100, and $k$ is the number of molecules having escaped through the hole. $W$ is the letter we will further use to denote the number of possible arrangements of our gas molecules for each situation (e.g. 0 escaped & 100 remain, $W(0)$, or 4 escaped & 96 remain, $W(4)$, et cetera).

    Here is a table with a few results. We included the situation where no single molecule has left the box. Obviously, the number of possible arrangements of the molecules leading to this situation, i.e. 0 escaped & 100 remain, is exactly one. We also included two more configurations where three and four particles have left the box. Note how quickly the possible arrangements increase.

    [table id=1 /]

    Probabilities

    How far can we take this? In this simplistic model, we can imagine the number of molecules remaining in the box becoming equal to the number having escaped into the other box: 50 escaped, 50 remain. So, let’s add $W(50)$ to the table. Also, let’s add more configurations to the table and have even more molecules escape the box until none are left, just to see what happens to the number of possible arrangements.

    [table id=2 /]

    As you can see, the number of possible arrangements decreases again after the box has reached its natural state of equilibrium (i.e. 50 escaped, 50 remain). This is, of course, only logical as the situation ‘flips’, as it were. More and more molecules end up escaping the box rather than remaining.

    If we were to calculate the probability of one or the other situations occurring, how should we go about this? Well, for instance, take the situation, or state, in which precisely zero gas molecules escaped. What is the probability of this state occurring?

    We would need to know the number of possible arrangements ($W$) in the state where there are 0 which escaped and 100 remain (that number is exactly 1, so, $W(0) = 1$), divided by the total of all possible arrangements in all states (the sum of all $W$’s). Mathematically, what we are calculating is the possibility $P$ where 0 molecules have escaped, in other words, for $P(0)$, we write:

    \begin{equation} P(0) = \frac{W(0)}{\text{total of all }W} = \frac{1}{\text{total of all }W}. \end{equation}

    Of course, the total of all $W$ still needs calculation. To do that, we use equation (1) to write down an equation for the sum of all $W$ in the previous table:

    \begin{align} \text{total of all }W &= \text{the sum of }W(0)\text{ to }W(100), \\ &= \sum_{n=0}^{100} \frac{100!}{n!(100-n)!}. \end{align}

    The answer to equation (4) is 1 267 650 600 228 229 401 496 703 205 376.

    That is a large number of total possible arrangements of all possible states. So, you can imagine that the probability of $P(0)$ occurring is inconceivably small: following equation (2), we get about $8 \times 10^{-29}$%. This would be 0% rounded to the nearest whole percentage.

    Likewise, we can calculate the probability of 50 escaping, 50 remaining, or P(50). This turns out to have a probability of 8% (to the nearest whole percentage). In the following plot of probabilities for each state, you can see which state is most likely to occur.

    A diagram showing that the equilibrium state (50 escaped, 50 remain) has the highest probability of just shy of 8%. Any other state has a drastically lower probability of occurring.

    So, in a way, given enough time, the arrangements of water molecules will converge to the state with the highest probability. Here, this is the situation where 50 molecules escaped, 50 remain. This is the so-called equilibrium point. Though there may be fluctuation around its equilibriumโ€”plus or minus one or a few particlesโ€”it is perfectly fair to say that the probability of no molecules remaining, P(0), and the probability of all molecules escaping, P(100), are near zero. Intuitively, this is what you would expect: while it is theoretically possible, in practice, you will never live long enough to ever witness all the molecules randomly gathering in just one box.

    Back to our wet clothes

    In real life, however, our pair of swimming trunks are not a box nor are there only one hundred water molecules. There are billions of water molecules in a liquid phase held together by the fabric of our garment. Also, there is the Sun. And there might be wind or even just a slight breeze. Besides, not looking like a box, swimmers don’t have a hole attached to a second box.

    However, if the box is a metaphor for our wet clothes, then the second box is a metaphor for its environment, the open air. And now, it gets interesting.

    In our example of the two boxes, no further external forces played any role(beginfootnote)A system that is thermally isolated from its surroundings is called an adiabatic system.(endfootnote). There was no wind, no Sun, no air humidity to consider. In reality, of course, they should be taken into account. Here’s what they do: Sun heats up the water in the trunks, causing its molecules to gain energy, aiding escape from the fabric. Wind causes the air molecules to bump into water molecules, removing excess water vapour around the clothes, aiding the water molecules to evaporate even further. As long as the relative air humidity isn’t too high, drier air helps to evaporate the water even further.

    So, what do these circumstances do, exactly? They shift the equilibrium point in our plot to the right. States where more and more molecules escape and less remain in the box, that is, stay in the swimming trunks, get a higher probability of occurring due to these circumstances.

    In other words, if our boxes would be subjected to the elements, it would cause the equilibrium point to shift from 50 escaped, 50 remain, towards 95 escaped, 5 remain, for instance.

    Ludwig Boltzmann (1844-1906)
    Ludwig Boltzmann (1844-1906)

    Moreover, since the air outside is practically infinitely large compared to our swimming trunksโ€”and not at all like the spatially limited second box of our metaphorโ€”the number of ways water molecules can be arranged by random motion, in a system of clothes hanging in the air outside, is significantly leaning towards the state where most escape into the air, even without wind and sunshine. Even though it may take longer, that state is practically inevitable.

    It was the Austrian physicist Ludwig Boltzmann who elaborated on this very statistical nature of states of a system, specifically in terms of its possible configurations of microscopically small molecules per one and the same end state.

    Entropy and the second law of thermodynamics

    Grave of Ludwig Boltzmann on Zentralfriedhof (Central Cemetery), Vienna, Austria

    Now, because the number of possible arrangements, $W(k)$, gets very big very fast, Boltzmann calculated their natural logarithm value. This is a very neat function when dealing with incredibly large numbers and exponential growth. On any respectable high school calculator, this can be done using the ‘ln’-button.

    Boltzmann then proceeded to multiply these log values with a constant $k$(beginfootnote)which is a different $k$ than the one we used earlier. The value of this $k = 1.380649 \times 10^{-23}\text{JK}^{-1}$.(endfootnote) to link the phenomenon of mechanically mixing stuff (arrangements of molecules) with the thermodynamical phenomenon of entropy of heat. We now call this constant $k$ the Boltzmann constant. He then wrote down the famous expression which we now call Boltzmann’s equation for entropy. In Vienna, in the city’s Central Cemetery, his gravestone is engraved with this very formula:

    \begin{equation} S = k\log_e W. \end{equation}

    So, now we get a new table with values for entropy $S$:

    [table id=3 /]

    As you can see, entropy $S$ increases towards the equilibrium point, only to decrease beyond that, up to the point where it is zero again. Note that the highest value of entropy also has the highest probability value. This means that the state of the box and its surroundings (the other box) will tend to maximum entropy. This also means that an equilibrium point entails maximum entropy.

    Going back to our system of wet clothing and their surroundings (the air) this means that, here too, the state of wet clothing will tend to maximum entropy. This value will correspond to the situation where most of the water molecules have escaped the clothes.

    This is the second law of thermodynamics: the entropy of the Universe tends to a maximum.

    Why do wet clothes dry?

    While external factors such as sunshine, wind, and relatively low air humidity do cause the probability distribution to shift more towards the state where most water molecules escape the clothes, based on the random motion of molecules alone, statistically, they should leave the fabric anyway (even though this is a slower process than with sunshine, wind, et cetera).

    This is because the number of ways in which water molecules remain inside the clothes is simply almost infinitely small compared to the number of ways where they are not inside the clothes. This leads to the statistical fact that the probability of water molecules not remaining in the clothing outweighs the probability of them remaining in the clothing.

    There are simply more places for water molecules to be in the open air than there are within the constrained spatial dimensions of someone’s tight swimming trunks. Or anyone’s, really.

    Ultimately, wet clothes dry because the entropy of the Universe tends to a maximum.


    Photo of Boltzmann’s grave by Daderot, under CC BY-SA 3.0.

  • Deriving the volume of the inside of a sphere using spherical coordinates

    Deriving the volume of the inside of a sphere using spherical coordinates


    Even though the well-known Archimedes has derived the formula for the inside of a sphere long before we were born, its derivation obtained through the use of spherical coordinates and a volume integral is not often seen in undergraduate textbooks.

    In this post, we will derive the following formula for the volume of a ball:

    \begin{equation}
    V = \frac{4}{3}\pi r^3,
    \end{equation}

    where $r$ is the radius.

    Note the use of the word ball as opposed to sphere; the latter denotes the infinitely thin shell, or, surface, of a perfectly round geometrical object in three-dimensional space. A surface has no volume, hence, we prefer to refer to it as a ball.

    This could be seen as a second-year university-level post.


    Spherical coordinates

    The volume of a cuboid $\delta V$ with length $a$, width $b$, height $c$ is given by $\delta V = a \times b \times c$.

    Figure 1: A volume element of a ball

    In Figure 1, you see a sketch of a volume element of a ball. Although its edges are curved, to calculate its volume, here too, we can use

    \begin{equation}
    \delta V \approx a \times b \times c,
    \end{equation}

    even though it is only an approximation.

    To use spherical coordinates, we can define $a$, $b$, and $c$ as follows:
    \begin{align}
    a &= PQ\delta\phi = r\sin\theta \, \delta\phi, \\
    b &= r\delta\theta, \\
    c &= \delta r.
    \end{align}

    So, equation (2) becomes

    \begin{align}
    \delta V &\approx r\sin\theta \, \delta\phi \times r\delta\theta \times \delta r, \nonumber \\
    &\approx r^2\sin\theta \, \delta\phi \, \delta\theta \, \delta r.
    \end{align}

    Volume integral

    Note that the relation becomes more precise when $\delta\phi$, $\delta\theta$, and $\delta r$ tend to zero. So, we can now write the volume integral for our ball $B$ as follows:

    \begin{equation*}
    V_B = \int_B dV_B = \int_\phi \int_\theta \int_r r^2\sin\theta \, dr \, d\theta \, d\phi.
    \end{equation*}

    Figure 2: To integrate over the infinite number of points (inside and on the surface) of a ball, one angle varies from $0$ to $2\pi$, which is $\phi$, in this case. Angle $\theta$ only needs to vary half of that as a ball is rotationally symmetric. Of course, bound by radius $r$.

    To set the upper and lower bounds for our integrals, we note that a ball has rotational symmetry about the $z$-axis (besides infinitely many others through the centre too). We will exploit this. We refer to Figure 2.

    Firstly, to integrate over infinitely many points between $0$ and $r$, the lower bound is $0$ and the upper bound is $r$:

    \begin{equation*} V_B = \int_B dV_B = \int_\phi \int_\theta \int_{r=0}^r r^2\sin\theta \, dr \, d\theta \, d\phi.
    \end{equation*}

    Secondly, to integrate over infinitely many points in the plane of angle $\theta$, we only need to regard the angles between $0$ and $\pi$,

    \begin{equation*}
    V_B = \int_B dV_B = \int_\phi \int_{\theta=0}^{\theta=\pi} \int_{r=0}^r r^2\sin\theta \, dr \, d\theta \, d\phi,
    \end{equation*}

    as we will proceed to, thirdly, rotate this plane, as it were, about the $z$-axis to integrate over infinitely many planes about said axis, which complete the shape of our ball. Hence, $\phi$ varies between $0$ and $2\pi$.

    And so, we calculate

    \begin{align}
    V_B = \int_B dV_B &= \int_{\phi=0}^{\phi=2\pi} \int_{\theta=0}^{\theta=\pi} \int_{r=0}^r r^2\sin\theta \, dr \, d\theta \, d\phi, \\
    &= \int_{\phi=0}^{\phi=2\pi} \int_{\theta=0}^{\theta=\pi} \left(\frac{1}{3} r^3\sin\theta \Big|_0^r\right) d\theta \, d\phi, \nonumber \\
    &= \frac{1}{3} \int_{\phi=0}^{\phi=2\pi} \int_{\theta=0}^{\theta=\pi} r^3\sin\theta \, d\theta \, d\phi, \nonumber \\
    &= -\frac{1}{3} \int_{\phi=0}^{\phi=2\pi} \left( r^3\cos\theta \Big|_0^{\pi} \right) \, d\phi, \nonumber \\
    &= \frac{2}{3} \int_{\phi=0}^{\phi=2\pi} r^3 \, d\phi, \nonumber \\
    &= \frac{2}{3} \left( \phi r^3 \Big|_0^{2\pi} \right), \nonumber \\
    &= \frac{4}{3}\pi r^3,
    \end{align}

    which is the desired result equal to equation (1).

  • Just a minute: what is a black hole?

    Just a minute: what is a black hole?


    Much like any question in the vain of โ€˜what is (…love…)โ€™, for which mathematical, physical, molecular, biological, psychological, philosophical, literary, and artistic approaches could be employed, here too, are several ways to approximate the answer to the question โ€˜What is a black hole?โ€™


    Letโ€™s take the notion of escape velocity, which is the minimum speed (in a specific direction) needed for an object to break loose of the gravitational domination of a massive body. Larger gravity means that you have to fly faster to get off the planet.

    Instead of a planet, letโ€™s pretend we have a rocket on the surface of the Sun. Why the Sun, you might ask. Rest assured, weโ€™ll definitely get to that. In the sketch below you see the situation at hand.

    Turns out, the larger the Sun’s mass, the stronger its gravitational influence on the rocket on its surface. Sounds obvious enough, right? In the sketch, the Sunโ€™s mass is symbolised by โ€˜big $M$โ€™. Of course, the rocket has a mass too, so this is denoted by โ€˜small mโ€™. Lastly, the distance between the centre of the Sun(โ€™s mass) and that of the rocket plays a big role. The larger the distance, the weaker the gravitational influence. This distance is denoted by the letter $r$. This obviously also means, the smaller the distance, the stronger the gravitational influence.

    Assuming the mass of the rocket (โ€˜small mโ€™) stays constant, we say that the gravitational influence is proportional to $M$ and inversely proportional to $r$. We can now begin to describe a mathematical relationship between the gravitational influence, $M$, and $r$. We can write:

    \begin{equation*} \text{gravitational influence} \propto \frac{M}{r}. \end{equation*}

    This weird $\propto$-sign means โ€˜is proportional toโ€™. The fraction $\frac{M}{r}$ means that, if $M$ grows bigger, the division grows bigger. If $r$ grows bigger, the division shrinks smaller. Letโ€™s just fill in some numbers to see how this works. Suppose, $M = 600$ and $r = 5$.

    \begin{equation*} \text{gravitational influence} \propto \frac{600}{5} = 120. \end{equation*}

    Letโ€™s make $M$ six times bigger: $M = 3600$. We get

    \begin{equation*} \text{gravitational influence} \propto \frac{3600}{5} = 720. \end{equation*}

    Not surprisingly, the division becomes six times larger too. If we make $r$ smaller, say $r=2$, the result becomes even larger:

    \begin{equation*} \text{gravitational influence} \propto \frac{3600}{2} = 1800. \end{equation*}

    So, what would happen ifโ€”in a thought experimentโ€”we would add more mass $M$ to our Sun? Indeed, the gravitational influence on the rocket would become larger. In turn, the rocket would have to fly faster in order to leave the Sun.

    What would happen ifโ€”in a further thought experimentโ€”we would not just add more mass $M$ to our Sun but also shrink its radius $r$, so that it becomes a small, very dense ball of stuff? Indeed, the gravitational influence would become even larger, so, the rocket would have to fly even faster.

    Schwarzschild

    The real formula that Newton came up with for the gravitational influence, which we call Newtonโ€™s law of universal gravitation, is a little different from what we have used up until now, and goes as follows:

    Of course, Einstein came up with an even more accurate set of formulas, but for our purpose, we wonโ€™t be using them as Newtonโ€™s law works just as well, in this case.

    John Michell, an 18th-century, English philosopher and clergyman, basically wondered if the ratio between mass $M$ and distance $r$ could lead to a gravitational influence so big that the required speed for a rocket to fly off to the stars would exceed the speed of light. He wasnโ€™t sure if anything with that much mass and such short a radius could ever exist, but if so, then, in theory, โ€˜dark starsโ€™ could exist.

    When some stars reach the end of their lives, they become supernovae. They explode their outer shells into space, while the inner shells of matter move inwards. This means that the surface quickly shrinks towards the centre of mass. This means that things, such as rockets, can get closer to the centre of mass while experiencing the gravitational influence of all that mass underneath it.

    Given the amount of the starโ€™s imploding mass $M$, at some point, there will be a distance $r$ around it where the gravitational influence is so large, that the minimum speed for a rocket to escape it will have to be larger than the speed of light.

    Karl Schwarzschild

    It was Karl Schwarzschild, a German physicist who found this distance using Albert Einsteinโ€™s equations of general relativity (the more accurate set of formulas compared to Newtonโ€™s).

    Given a certain mass, the Schwarzschild radius is the distance from the centre of mass below which the magnitude of the escape velocity is larger than the speed of light. And so, this part of the universe will be black, whence no one returns. Every point around the centre of this mass as described by the Schwarzschild radius, forms what we call the event horizon.

    A black hole is thus a region in the universe where the gravitational influence is so large that nothing, not even light, can escape it. Or, to put it a bit more technically, it is a region of spacetime where every possible future leads to its singularity. We might explain the latter in another article, in the future.

    Event Horizon Telescope

    A vast array of radio observatories and telescope facilities around the world basically turned our entire planet into one big telescope, called the Event Horizon Telescope.

    On Wednesday, 10 April 2019, at 15:00 CET, the first photo of a supermassive black hole in the middle of a galaxy called Messier 87 was presented. Later, a photo of the black hole in our own Milky Way, called Sagittarius A*, will be expected.

    The observations may test Einsteinโ€™s general relativity yet again. Perhaps more on that in another article.

    We highly recommend watching the recording of the live stream of the presentation of the results.

    We have been focussing on non-rotating black holes. The physical models for a rotating black holes differ to some degree, but not significantly for the scope of this article.

    Featured image: the image of the supermassive black hole Messier 87. Credit: EHT Collaboration

  • Simple problems on relativistic energy and momentum

    Simple problems on relativistic energy and momentum


    We will focus on a few simple problems where we will manipulate the equations for relativistic energy and momentum.

    This could be seen as a second-year university-level post.


    Einstein had shown that the Lorentz transformations were the correct way to switch between the coordinate systems of different frames of reference [1]. He also taught us that Newtonโ€™s laws werenโ€™t at all proper relativistic laws. For instance, Newtonian momentum $ \mathbf{p} = m \mathbf{v} $, and energy $ E = mv^2 / 2 $ were not at all accurate at speeds approaching that of light.

    Instead, we have all come to learn that the relativistic momentum is written as

    \begin{equation} \label{eq:relativistic momentum} \mathbf{p} = \frac{m \mathbf{v}}{\sqrt{1 – \dfrac{v^2}{c^2}}}. \end{equation}

    And that the correct relativistic expression for total energy is

    \begin{equation} \label{eq:relativistic energy} E_{\text{tot}} = \frac{mc^2}{\sqrt{1 – \dfrac{v^2}{c^2}}}. \end{equation}

    We will solve the following problem set:

    1. Prove, for a particle travelling at $ c $, that the magnitude of the relativistic energy is given by $ E = pc $.
    2. Show that the energy-momentum relation for a particle with any mass $ m $ travelling at any speed $ v $ is correct and do mind it is not the famous $ E = mc^2 $ we are referring to. Use the correct one, if you please.
    3. Given that the mass of a proton is $ m_p $, calculate its exact speed when its relativistic  translational kinetic energy (which is the relativistic total energy minus its relativistic mass energy) is four times its relativistic mass energy.

    Problem I

    Since $ E $ is expressed in terms of $ p $, we need to rewrite Eq. $ \eqref{eq:relativistic momentum} $ by solving for $ m $:

    \[ m = \frac{p \sqrt{1 – \dfrac{v^2}{c^2}}}{v}. \]

    Note, we do not use the vector quantities, just the magnitudes. We can now proceed to substitute this into Eq. $ \eqref{eq:relativistic energy} $:

    \[ E_{\text{tot}} = \frac{\left(\dfrac{p \sqrt{1 – \dfrac{v^2}{c^2}}}{v}\right)c^2}{\sqrt{1-\dfrac{v^2}{c^2}}}. \]

    This reduces to

    \begin{align}
    E_{\text{tot}} &= \frac{pc^2 \sqrt{1 – \dfrac{v^2}{c^2}}}{v \sqrt{1 – \dfrac{v^2}{c^2}}}, \\
    \therefore E_{\text{tot}} &= \frac{pc^2}{v}. \label{eq:E=pc^2/v}
    \end{align}

    As we are dealing with a particle travelling at speed $ c $, we know $ v = c $, rendering Eq. $ \eqref{eq:E=pc^2/v} $ to

    \begin{align}
    E_{\text{tot}} &= \frac{pc^2}{c}, \\
    \therefore E_{\text{tot}} &= pc.
    \end{align}

    Problem II

    The energy-momentum relation is

    \[ E^2_{\text{tot}} = p^2c^2 + m^2c^4. \]

    Substituting Eqs. $ \eqref{eq:relativistic momentum} $ and $ \eqref{eq:relativistic energy} $, yields

    \[ \left(\frac{mc^2}{\sqrt{1 – \dfrac{v^2}{c^2}}}\right)^2 = \left(\frac{m \mathbf{v}}{\sqrt{1 – \dfrac{v^2}{c^2}}}\right)^2c^2 + m^2c^4, \]

    which we can continue to work out as follows:

    \begin{align*}\left(\frac{mc^2}{\sqrt{1 – \dfrac{v^2}{c^2}}}\right)^2 – \left(\frac{m \mathbf{v}}{\sqrt{1 – \dfrac{v^2}{c^2}}}\right)^2c^2 – m^2c^4 &= 0, \\
    \frac{m^2c^4}{1 – \dfrac{v^2}{c^2}} – \frac{m^2v^2c^2}{1 – \dfrac{v^2}{c^2}} – m^2c^4 &= 0, \\
    \left(1-\dfrac{v^2}{c^2}\right)\left(\frac{m^2c^4}{1-\dfrac{v^2}{c^2}}\right) \qquad &\qquad \\ – \left(1-\dfrac{v^2}{c^2}\right)\left(\frac{m^2v^2c^2}{1-\dfrac{v^2}{c^2}}\right) &\qquad \\ – \left(1-\dfrac{v^2}{c^2}\right)m^2c^4 &= 0, \\
    m^2c^4 – m^2v^2c^2 – m^2c^4 + \frac{m^2v^2c^4}{c^2} &= 0, \\
    m^2c^4 – m^2c^4 – m^2v^2c^2 + m^2v^2c^2 &= 0, \\
    0 – 0 &= 0.
    \end{align*}

    Hence, for every value of $ m $, $ p $, and thus $ v $, the relation holds.

    Problem III

    Hydrogen bubble chamber Fermilab

    The relativistic (total) energy is

    \[ E_{\text{tot}} = E_{\text{trans}} + E_{\text{mass}}. \]

    If the relativistic translational kinetic energy is four times the relativistic mass energy, then we can write

    \[ E_{\text{trans}} = 4E_{\text{mass}}. \]

    In our case, this then yields for the relativistic (total) energy:

    \[ E_{\text{tot}} = 4E_{\text{mass}} + E_{\text{mass}} = 5E_{text{mass}}. \]

    To calculate the protonโ€™s speed, we then write

    \begin{align*}
    \frac{m_pc^2}{\sqrt{1 – \dfrac{v^2}{c^2}}} &= 5E_{\text{mass}} = 5m_pc^2, \\
    \frac{1}{\sqrt{1 – \dfrac{v^2}{c^2}}} &= 5, \\
    \sqrt{1-v^2/c^2} &= \frac{1}{5}, \\
    1-\frac{v^2}{c^2} &= \frac{1}{25}, \\
    \frac{v^2}{c^2} &= \frac{24}{25}, \\
    v^2 &= \frac{24c^2}{25}, \\
    \therefore v &= \sqrt{\frac{24c}{25}} = \frac{2\sqrt{6}c}{5},
    \end{align*}

    which is about $ 0.98c $ rounded to two decimals, which means that the proton zips at about 98% of the speed of light through the fabric of the cosmos.

    Image Hydrogen bubble chamber Fermilab: Proton with 300 GeV energy producing 26 charged particles in the 30 inch hydrogen bubble chamber at Fermilab. Source: Wikimedia Commons

    [1] Einstein, A. (1905) โ€˜Zur Elektrodynamik bewegter Kรถrperโ€™, Annalen der Physik, 322(10), pp. 891โ€“921. doi: 10.1002/andp.19053221004.


    This is a repost. Slight errors in the parsing of LaTeX in the original article of 24 December 2018 have been corrected.

  • Just a minute: how big is the universe?

    Just a minute: how big is the universe?


    In case a (your) child asks how big the universe is exactly, you might want to know what the most honest, straightforward answer is. Although there is, as of yet, no feasible way of getting to know its size with certainty, we do have some cool things such as telescopes and mathematics at our disposal. When it comes to estimating the size of what we see when we look up, they help a lot.

    Observable universe

    Imagine yourself floating in pitch darkness. You have no knowledge of where you are. Fortunately, you have binoculars in the left pocket of your spacesuit. As your eyes become adjusted to the lenses, you start to distinguish a faint point of light. Then you remember, you have an even better extensible monocular telescope in your other pocket. Like a pirate in the night, you gaze through the glass. And then you discover, it wasn’t just a dot of light, it was a whole group of dots of light.

    We see a person dressed in an astronaut suit floating in space. It is looking at a giant spiral galaxy floating in the same space.

    You decide to swirl the telescope around and, since you’re afloat in the middle of, well, nothing, or so it seems, you can even look down, underneath your feet. As you scan all around, you discover there is a nearly uncountable number of groups of dots of light all around you, like a sphere of clustered dots with you at its centre.

    Valuable lessons from science class taught you that light has a certain maximum speed, so you realise that, maybe, more groups of dots of light are there, beyond the ones you see, but that their light hasn’t reached you yet.

    The part of your universe you can in principle seeโ€”limited only by the speed of light and not whether or not we have the technologyโ€”is what we call the observable universe. In fact, this is the type of thing we have the Hubble Telescope do for us.

    Now, here’s the problem: the universe might just be way bigger than the observable part, but we won’t know by how much because it is highly unlikely light from beyond what is now observable will ever reach us. This is because the universe is expanding rapidly, and more rapidly every second.

    The yellow sphere represents the observable universe with Earth at its centre. Everything outside its radius is purely hypothetical. Whatever the location of Earthโ€”more downward, more to the leftโ€”Earth will remain the centre of its observable region.

    About you being at the centre, think of yourself floating around in space once more. If you look around you, the separation between you and every point on the sphere of your observation is equidistant. And if you were to float to a different region, your observational sphere moves along with you. So, you are always at the centre of your observation bubble.

    The same principle is valid for our place in the observable universe, on Earth. And I’m not talking about social-psychological information bubbles on the internet, I mean literally, physically, factually. Mind you, this is not the same as saying we are at the centre of the universe as there is no centre of the universe. You are, however, always at the centre of the observable universe. Therein lies the difference.

    Accelerated expansion

    Light from beyond the observable part cannot reach us because the universe is growing orโ€”as cosmologists prefer to sayโ€”expanding. And it is doing so at an ever faster-going rate. Per second, more space is added than light can cover.

    Imagine you’re inside a dream, standing in the narrow hallway of a hotel. You decide to walk to the door at the other end. Now imagine the hallway suddenly growing longer and longer. The walls on either side are not just being optically elongated, but structural bits of the wall seamlessly seem to multiply physically. The hotel is not just stretching, it is expanding its material. In the meantime, you’re moving, but you’re not moving forward by much. The rate of the hallway’s expansion is greater than your rate of displacement.

    We see a girl or woman walking down a hallway. While she is walking, the hallway seems to stretch out, or expand. The door at the other end recedes from the walking person.
    This hallway is a metaphor for the expanding universe. In reality, the camera operators of the movie Poltergeist (1982) applied the famous vertigo effect, an illusion causing the walls to stretch optically. In our metaphor, however, the ‘stuff’ of the walls is being multiplied seamlessly.

    Well, this is what a lonely photon must be experiencing, emanated from a distant star, on its way to us, but never reaching us since the space it is travelling through is expanding. Note that the door in the GIF above is not moving, it is spaceโ€”the hotel structureโ€”that is expanding! The door is a metaphor for distant galaxies. It’s not the galaxies that recede, it is space expanding.

    Estimated size

    At the current expansion rate, and given the age of the universe, the observable universe is estimated to be 93 billion light years in diameter. Which is about 880 000 000 000 000 000 000 000 kilometre (546 800 000 000 000 000 000 000 miles).

    This means that its volume is about 400 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 litres ($ 4 \times 10^{83} $ l) or a 9 followed by 82 zeros in gallons.

    Earth has a volume of about 108 300 000 000 000 000 000 000 litres. So, it occupies about 0.000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 003% of the observable universe.

    The size of the universe, however, is unknown. Could be finite, could be infinite. At present, there is no reason to accept either statements definitively.

    ESO animation

    Have a look at this cool animation made by the European Southern Observatory (ESO) and others. We fly outward from the ESO Supernova Planetarium & Visitor Centre at ESO’s headquarters in Garching bei Mรผnchen to the place where we see a vast number of galaxies. Realise that, as soon as the video takes us out of our own galaxy, which itself is starting to look like ‘just another star’, every other ‘star’ you start seeing is, in fact, an entire galaxy of its own. Of course, the film ends with a fade-out for we do not know whether there are any borders or if it goes on forever.

    Screenshot of ESO’s animation. Click to open the YouTube clip.

    Epilogue

    The attentive reader will have noticed we only talked about the size of the observable universe, yet, we did not mention what we see when we look at its very edgeโ€”as far as current technology permits.

    As it takes such a long time for that light to have reached us, the stuff that we see, is old. We are looking at the past, the history of the universe. A glance at the cosmos is like time-travelling back to days of old.

    This deserves a separate article. Unfortunately, we do not have space nor time right now.


    Photo Observable universe adapted on the basis of Strogoff‘s work published under CC BY-SA 3.0.

  • Professor Karen Uhlenbeck wins the prestigious Abel Prize 2019

    Professor Karen Uhlenbeck wins the prestigious Abel Prize 2019


    Karen Keskulla Uhlenbeck received the prestigious Abel Prize 2019 for her revolutionary theories in geometric analysis and gauge theory. Her mathematical work proved to be fundamental in our understanding of minimal surfaces in our three-dimensional space and beyond. Moreover, she played an essential role in laying the foundation for many modern mathematical models used in particle physics, string theory, and general relativity.

    Currently, Professor Emerita of Mathematics and Sid W. Richardson Regents Chair at the University of Texas at Austin, Uhlenbeck is a Visitor in the School of Mathematics at the Institute for Advanced Studies, Princeton.

    “Quite frankly: it is about time.”
    โ€”Prof Helmut Hofer (IAS)

    Naturally, this momentous event at the Norwegian Academy of Science and Letters has been preceded by many extraordinary highlights in her career as a highly sought-after pioneering expert. Some of them stand out even more because they expose the perplexing disparity in the acknowledgement of contributions between men and women in academia. In 1990, she held a Plenary Lecture at the world’s most seminal International Congress of Mathematicians (ICM), in Kyoto, making her only the second woman in history to do so. The other one being the eminent Emmy Noether forty-eight years earlier.

    Today brought another highlight: she is the first woman to receive the Abel Prize.

    “Quite frankly: it is about time. Karen has had a tremendous impact on the development of modern geometric analysis, particularly the calculus of variations. Her contributions to minimal surface theory and Yang-Mills theory have changed the subjects and started some of the most exciting developments in mathematics,” said Helmut Hofer, IAS Professor in the School of Mathematics, in the Institute’s press release.

    For more information about Professor Uhlenbeck, her achievements, and on the Abel Prize, visit their website.

    Featured photograph: In 1987 Karen Keskulla Uhlenbeck moved to the University of Texas at Austin to take up the Sid W. Richardson Foundation Regents’ Chair in mathematics where she worked until 2014. Currently, Uhlenbeck is a Visiting Senior Research Scholar at Princeton University as well as Visitor in the School of Mathematics at the Institute for Advanced Study (IAS). Photo: Andrea Kane/Institute for Advanced Study

  • Happy birthday mister Einstein, happy Pi Day to you!

    Happy birthday mister Einstein, happy Pi Day to you!


    ฮ  Day is the day on which we commemorate Albert Einsteinโ€™s (1879-1955) birthday. Also, people celebrate the existence of $ \pi $ as today is 3/14, forming the first three digits (at least) of the number $ \pi $ in the American date format. Some Western European criticsโ€”on Twitter, for exampleโ€”have stated one oughtnโ€™t as โ€˜we, hereโ€™ simply do not use the American date format. Of course, nearly the whole rest of the world do not use the American date formatโ€”hence, โ€˜Americanโ€™โ€”but it hasnโ€™t stopped cheerful people from all over that same rest of the world to celebrate and put mathematics into the limelight once a year.


    Larry Shaw (1939-2017), the founder of Pi Day, at the Exploratorium in San Francisco

    In 1987 or 1988, a physicist named Larry Shaw (1939-2017), while working at the Exploratorium, museum for science, art, and human perception, came up with the idea of celebrating the mathematical constants on March 14th. What started out as eating pie with just his colleagues, the event became public the next year. At 1:59pm, a time notation predominantly used in the US and the Commonwealth, forming (at least) the fourth, fifth, and sixth digits, a parade would be held with each visitor holding a digit of pi while eating pie and singing happy birthday to Albert Einstein. Larry was pleased to see the younger visitors loving the museumโ€™s festivities, which, furthermore, include pi poetry readings, pi-kus (haikus about pi) and pi limericks, a pizza-dough tossing lesson, and eating it.

    Hidden pis

    (Grow up, itโ€™s not even spelt right.) One of the most fascinating things about pi is that it tends to come up in places where you would least expect it. For instance, Albert Einstein and pi have a relationship. His general theory of relativity pivots around the following field equations:

    \[ R_{\mu\nu}-\frac{1}{2}Rg_{\mu\nu}=8\pi GT_{\mu\nu}. \]

    We wonโ€™t get into the details, but itโ€™s pretty delightful that a theory describing one of the most fundamental forces in our universe, called gravity, would need the ever so humble pi.

    A long string of digits has been incorporated into the calรงada portuguesa thanks to mathematics teacher and current chair of Faro’s city council Rogรฉrio Bacalhau. Credits: @kjrunia, licensed under CC BY 4.0.

    And this one is even cooler. Mathematicians wondered what you would get when you sum the following series of terms to infinity:

    \[ \frac{1}{1^2}+\frac{1}{2^2}+\frac{1}{3^2}+\frac{1}{4^2}+\dots \]

    The genius mathematician Leonhard Euler solved this Basel problem and found that the sum would converge to $ \pi^2/6 $. Even when a series tends to infinity, the ever so humble pi appears.

    Speaking of ‘humble pi’, recently, a great book with this very title has come out by my favourite stand-up mathematician and YouTuber Matt Parker. I recommend it. Itโ€™s great. In this video, he is trying to approximate pi by using classical mechanics. Do have a look! Over the years, he made a whole bunch of cool and funny videos calculating pi. If you find yourself trapped in the algorithmic funnel that its inventors called YouTube, you’re welcome.

    Screenshot of Matt Parker’s YouTube video in which he is calculating pi using a balancing beam.

    One of the most fascinating places where pi pops up is where billiard balls bounce against each other and the cushion on the inner rail of a billiard table. Gregory Galperin at the Department of Mathematics of the Eastern Illinois University wrote a paper demonstrating how pi could be obtained in a jaw-droppingly awesome way.

    The New York Times published a blog post about it in 2014 but not before the YouTube channel Numberphileโ€”another favouriteโ€”had professor Ed Copeland explain it already in 2012.

    Recently, however, the YouTube channel 3Blue1Brown published a video about it too. (Yes, the channel is also a favourite and I realise that I am using the word in a contradictory manner.)

    It features a gorgeous simulation and is somehow very pleasing to the ears. Also, Grant Sanderson, the mathematician behind the voice and videos, does a great job of visually deciphering the language of the universe. Do have a look. He then gives the answer as to ‘but how’ and ‘why at all’ in a second video.

    If you haven’t seen it, do support your chin firmly with your hand while letting the video play out as it may gravitate towards the centre of Earth, radially.

    A screenshot of 3Blue1Brown’s video on calculating pi using collisions.

    Photo of Larry Shaw: credits: Ronhip, licensed under CC BY-SA 3.0.
    Photo of digits of pi in the Portuguese streets: credits: @kjrunia, licensed under CC BY 4.0.

  • The formula that got Albert Einstein the Nobel Prize and should stop us getting sunburn all the time

    The formula that got Albert Einstein the Nobel Prize and should stop us getting sunburn all the time


    A copy of page 5 of the newspaper The Times of 10 November 1922. Near the bottom, a small article is printed. The title is Nobel Prize for Einstein. The text goes as follows. Stockholm, Nov 9.โ€”The Nobel Prize for Physicsโ€”1921โ€”has been awarded to Professor Albert Einstein, of Berlin, in recognition of his work in theoretical physics. The 1922 prize for physics has been awarded to Professor Niels Bohr, of Copenhagen, in recognition of his research work into the structure of atoms.โ€”Reuter.
    ‘Nobel Prize for Einstein’, one sentence was spent in The Times of 10 November 1922.

    In 1921, Albert Einstein won the Nobel Prize “for his services to Theoretical Physics, and especially for his discovery of the law of the photoelectric effect.” Not a word about relativity. So, no, he did not win the Prize with $ E=mc^2 $. Though it is his most famous equationโ€”which, by the way, is not the complete versionโ€”it is not his Nobel Prize-winning formula. We will write it down, but first, we describe what this photoelectric effect is.


    Different stuff is made up of different molecules. Different molecules are made up of different atoms. Different atoms are made up of a variety of nuclear composites and different numbers of electrons. So far, nothing new, perhaps, but here’s the thing. If electrons are exposed to particular amounts of energy, they can be ejected away from the nucleus.

    An atom of which one or more electrons have been blasted away is called an ion. The process is called ionisation. Whether this occurs, depends on a few things such as the type of stuff (=the type of atoms and how they are bound together) and the specific energy it is exposed to.

    If ionisation at the surface of a material is achieved by normal light, we call this the photoelectric effect: light (the ‘photo’-part) causing electrons to leave their nucleus (the ‘electric’-part).

    A diagram of the ionisation of an atom (not to scale). (1) The yellow cloud represents an electronโ€™s (probable) whereabouts. The tiny pink core represents the atom’s nucleus. (2) Photons of a specific colour radiate towards the atom. (3) The electron has flown off. The nucleus remains. The atom has become an ion.

    Not about intensity

    One peculiar thing is worth mentioning. In fact, it was this puzzle that led Albert to his equation. It turned out that what matters is the frequency of the light beam, i.e. the colour of the light, not the intensity of it, i.e. the power per square metre, or Joule per second (watt) per square metre.

    Imagine, in the diagram above, that a billion yellow photons would radiate towards the atom and nothing happened; the electron would stay where it was. Now imagine a billion billion billion billion yellow photons approaching the atom. Still nothing would happen as it is not about intensity.

    Yellow light is less energetic than blue light, so if you would replace the light bulb for a source that delivers pure blue light, with even one blue photon, it could happen easily (though you would have to aim impossibly precise, so it makes sense to actually radiate a lot). This puzzled many scientists, but Albert solved it and won the Nobel Prize.

    With his discovery, quantum physics was starting to get momentum. He, and other good physicists of his time, showed that light could be seen as little packets of energy, which scientists started calling photons. A beam of light was now a stream of photons. The intensity, the amount of photons per second per square metres doesnโ€™t matter but the frequency of a photon, or energy per photon does.

    DNA

    While this is all cool and useful for scientific purposes, we certainly do not want any electrons of the DNA molecules of our skin breaking away from their atomic confines. Atomic bonds would be destroyed and our DNA would become mutated. Even though astonishing molecular biological processes in our body repair defects like this in a staggering, basically inconceivable number of cases, some errors might slip through and may even become the start of tumour growth. Therefore, it is important to know what energy domains would cause our beloved bodily electrons to be blasted off so that humanity can learn to avoid those dangerous environments.

    The problem arises when we get into the mid to high-energy electromagnetic radiation, or light, or photons, if you will. We’re talking the dangerous kind of ultraviolet here, the type of UV causing DNA mutation to occur: UVB to be precise. A photon of UVB-light is about 1.8 times more energetic than a photon of the yellowish light in your home and almost a million times more energetic than a mobile phone photon. So, don’t be scared of being home. As soon as you set foot outside, though, be afraid. Not of the dark, but of the light, for ionising UVB-light is emitted by the Sun.

    A diagram of electromagnetic radiation. Far right, we see the dangerous types of radiation: cosmic rays, x-rays, gamma rays, UV-light. In the middle, we see visible light. Far left, we see the lowest energy photons: WiFi, mobile phones, microwave ovens.
    A diagram (not to scale) of electromagnetic radiation, or photons, if you will. The mentioned values are the frequencies of the photons, expressed in gigahertz (GHz). The higher the frequency, the higher the energy of the photon.

    Fortunately, as stated before, our bodies have evolved to repair the damage when necessary. This is why even X-rays are okay and hospitals and dentists make sure not to expose you to doses of energetic photons you wouldn’t survive. Continuous monitoring of its uses and effects is prerequisite.

    It’s partly a question of the law of large numbers, though. If the number of freely whizzing electrons is large enough, they themselves will become the main cause of an increasing number of damaged DNA molecules, and, eventually, some repairs will fail or not even take place. So, while it is not instantly dangerous, we do recommend some reading up on the subject of sunbathing. Use UV protection. Don’t get sunburnt. And give your body a chance to recover from the ruthless blasts of ionising UV radiation. Forget microwaves, the problem is crispy skin.

    The formula

    So, now we finally get to Albert’s Nobel Prize-winning formula. Here it is

    \[ \frac{1}{2}m_ev^2_\text{max} = h\nu – \phi. \]

    It doesn’t look as sassy as the other one, right? And yet, it’s the one that allows us to calculate if, for instance, electrons of our body’s carbon atoms get blasted out by the photons emitted by the lamp in your lavatory (they do not). Or if the laser pointer knocks some electrons out (it doesn’t), which we use anyway, because we need to point at things on our PowerPoint slides as they might well be ill-designed (they are).

    So, $ \frac{1}{2}m_ev^2_\text{max} $ means maximum kinetic energy, which is simply the energy with which an electron flies away from its nucleus. If its value turns out to be smaller than or equal to zero then the electron is not affected at all. It’ll keep stuck to its nucleus. If it is larger than zero then off it goes. The symbol $ h $ is a constant, which we needn’t worry too much about. It’s a number and it’s called the Planck constant. The Greek letter $ \nu $ is the frequency of the photon. In the diagram above, a few have been mentioned. Mind you, $ h\nu $ means $ h \times \nu $ and is the energy of a photon. Mathematicians, physicists, engineers, and other folks, just like to leave out the $ \times $-sign. The Greek letter $ \phi $ is the so-called work function. It is the minimal energy needed for the occurrence of a photoelectric effect. Its value depends on the type of atom, molecule, material, and surface you want to calculate the photoelectric effect of.

    In conclusion

    Notice Einstein’s formula does not have any term relating to the number of photons radiated per second per square metre towards the atom of interest, i.e. the intensity. Only the frequency is important. This means that atomsโ€”such as your bodyโ€”will be left undisturbed irrespective of the power of the radiation they are exposed to. There may be a bit of heat but there is no ionisation. The potential danger lies in frequency ($ \nu $), such as that of UV light and higher. Here, both dosage and capability of recovery play a crucial role.

    Young Albert Einstein

    The value of the Planck constant is $ h = 6.626070 \times 10^{-34} $ Js (Joulesecond). The value of the work function of carbon, of which our entire body is made, including our DNA, is $ \phi = 8.0108831 \times 10^{-19} $ J. If a WiFi photon has a frequency of 2.5 GHz, you can calculate yourself if it would yank the electrons from a carbon atom. Remember to convert 2.5 GHz to $ 2.5 \times 10^9 $ / s (per second). Thanks to Albert, calculating this has become child’s play. We could do the maths on the back of an envelope. If all the terms on the right hand side of the equal sign turn out to be larger than zero, then sell your router immediately andโ€”based on this diagramโ€”you most definitely ought to refrain from switching on the light while frequenting the lavatory. Good luck with the calculation! (Or check the working out.)


    Featured image: a 14-year-old Albert Einstein, photographed in 1893. Credits EMILIO SEGRE VISUAL ARCHIVES / AMERICAN INSTITUTE OF PHYSICS / SCIENCE PHOTO LIBRARY / Universal Images Group. Source: Young Albert Einstein, physicist. [Photography]. Encyclopรฆdia Britannica ImageQuest. Retrieved 9 Mar 2019, from 
    https://quest.eb.com/search/132_1258083/1/132_1258083/cite

    Smaller image of an even younger Albert Einstein: Credits EMILIO SEGRE VISUAL ARCHIVES / AMERICAN INSTITUTE OF PHYSICS / SCIENCE PHOTO LIBRARY / Universal Images Group. Source: Young Albert Einstein, physicist. [Photography]. Encyclopรฆdia Britannica ImageQuest. Retrieved 9 Mar 2019, from https://quest.eb.com/search/132_1255429/1/132_1255429/cite

    Newspaper article: “Nobel Prize for Einstein.” Times, 10 Nov. 1922, p. 5. The Times Digital Archive. Retrieved 8 Mar 2019 from http://tinyurl.galegroup.com/tinyurl/9Q37o0.

  • Just a minute: why do large and heavy ships not sink?

    Just a minute: why do large and heavy ships not sink?


    Until they do due to a mistake, ships do not sink, not even the large and heavy ones. Now and then, textbooks say this is because of dissimilar density. Though not wrong, it is also not a fundamental reason. While ships may sink to the bottom of the ocean thanks to gravity, they also float thanks to gravity.


    When a vessel is launched in the water, it will always sink a little bit under the surface, until it stops sinking, preferably at a safe distance from where humans tend to loiter. And since, in this universe, water and the submerged part of a hull cannot occupy the same space at the same time, the submerged volume equals the volume of displaced water. This causes, however small, to raise the surface of the water. Due to the sheer size of most bodies of water, that rise is unnoticeable.

    In spite of this usually insignificant level increase, gravity is still ‘pulling down’ every cubic part of the raised water, which is then, through pressure, also pushing on the ship. The force of the weight of the displaced water is equal to the force exerted upwards on the bottom of the ship. This is called Archimedes’ principle.

    In other words, while the ship exerts a force on the water due to gravity, the water around the ship exerts a force back at it, through pressure, due to gravity. Notice how it is working against itself, as it were. But, as long as the force of the weight of the water is equal to the force of the weight of the ship, it’s fine. Yes, water pressure also pushes on all submerged sides of the object, but they cancel each other out as they work against each other with equal strength so we can leave them out of the equation.

    The trick, of course, is to design the shape of a hull in such a way that its submerged volume displaces a volume of water weighing as much as the ship’s weight. These choices influence the ratio between its volume and its mass. And the latter is why referrals to density are madeโ€”often accompanied by a nifty display of algebra. Though not fundamental, density is a useful property to work with on Earth, such as when explaining why oil floats on water.

    Until you are not on Earth but on the International Space Station, for instance. Do have a look at what happens when the lower-density oil and higher-density water are put together when gravity is out of the mix.

    In Figure (1), a mass is launched in the water. The water level is indicated by the dashed line. In Figure (2), part of the mass is submerged, thereby displacing upward a certain volume of water left and right. The grey arrow denotes the (force of the) weight of the mass. The downward blue arrows denote the (force of the) weight of the displaced volume of water. The latter two cause an upward pressure to the bottom of the mass, as denoted by two upward arrows. Notice how the sum of the length of these two arrows equals the length of the grey arrow: our mass is buoyant.

    Of course, air pressure also exerts a force on the ship. However, it does so on the water surface too. As we also wanted to keep things simple, we thus did not take this any further into consideration.

  • Mirror, mirror, what’s up with the mirror writing?

    Mirror, mirror, what’s up with the mirror writing?


    Ever wondered why sentences, words, and letters always exclusively seem to have their left and right reversed in the looking glass, while mirror writing is almost never projected upside down? Things are happening which may not be obvious. For starters, mirrors do not reverse left and right.

    We are intelligent types with (sometimes too much) self-awareness. We look in the mirror, and we know it is our reflection and not someone else staring back at us. However, if you were ever under the impression that, for instance, the left and right sides of your face are reversed, you might want to re-evaluate the depth of this appreciation. You are probably still mistaking your mirror image for a real other person facing you. Even though the situations look similar, they are not equalโ€”not just in the metaphysical sense but, more relevantly, in the mathematical sense. This relates to why letters, words, and sentences seem left-right-reversed by the mirror and almost never projected upside down, but we will get to that later.

    Mirror Guy

    Let’s reflect on the photo below for a moment. The Tie Guy in front of the mirror, trying to tie his tie for a white tie dinner, is looking at himself in the mirror. Who knew? I know but bear with me, because both peculiarly and crucially, it is important to acknowledge that it is, in fact, himself, and not someone else.

    Photography: Pete Souza

    Let’s perform a thought experiment. Imagine Tie Guy drawing a big L on the palm of his left hand and a big R on that of his right.

    1. Tie Guy has an L on the palm of his left hand and an R on the palm of his right hand;
    2. Mirror Guy is not someone else: Mirror Guy is Tie Guy;
    3. Tie Guy presses his left hand with an L against the mirror;
    4. if a mirror would reverse Tie Guy’s left and right, Mirror Guy should use the hand with an R, since that is Tie Guy’s right hand;
    5. Mirror Guy does not, he uses the hand with an L;
    6. hence, Tie Guy’s left hand with an L is Mirror Guy’s left hand with an L;
    7. therefore, a mirror does not reverse left and right.

    Point 6 is probably hardest to grasp at first. Clearly, the Ls of both Guys are up against the mirror. Yet, we still think a mirror reverses left and right. The cognitive hurdle is perhaps that while we accept our mirror image to be us, we are hardwired to continue to treat it as though it were someone else facing us. On a daily basis, chances are we interact more with people facing us than we do with our mirrored selves.

    We have grown accustomed to mentally reverse left and right, which evidently has proven to be useful during our interactions with everyday humans. We were taught already at an early age that ‘your left is her right’ or ‘her left is your right’. Similarly, a mathematical physics professor, upon turning around, causing her to face the audience in the lecture hall again, knows all too well that the strings of equations on the blackboard on her left are in fact on her studentsโ€™ right.

    Ironically, left and right get reversed in the real world and not in mirrors. Our left hand is simply our mirror image’s left hand and not ‘their right’โ€”that is just how the Mirror Universe works.

    Transformation

    Wait, did I just write ‘turning around’ in italics for a particular reason? Indeed, I did. Turns out, rotations, reflections, and symmetries are tricky. Hence, what follows is an important distinction.

    Mathematically, someone facing youโ€”be it your biological cloneโ€”has been rotated (180 degrees) with respect to your position and direction, whereas your mirrored self is not. Instead, your mirror image is a… well, a reflection. (Surprise.) Here is the crux: rotation and reflection are distinct geometrical transformations. Stating that a mirror reverses left and right is similar to confusing rotation with reflection.

    Rotation and reflection are not the same

    After a rotation of 180ยบ, we need to use the antonym of ‘left’ or ‘rightโ€™, which is not needed in the case of a reflection. However, it is highly plausible one does not encounter reflected human beings very often. We often deal with rotated bipeds. So, in the case of looking at your mirror image, you just need to un-think your mirrored self is someone else. In a way, even more than you would expect, being the conscious, intelligent life form that we are supposed to be, you need to accept that the mirrored self is you.

    Funnily enough, accepting that a mirror does not reverse up and down either is, without doubt, a lot easier. Imagine Tie Guy banging his head against the mirror. Mirror Guy does not then bump his feet against the mirror. Therefore, a mirror does not reverse up and down nor left and right.

    Why mirrored letters look weird

    So, what’s up with the mirror writing? Clearly, something is going on with letters, words, and stacks of writings as Leonardo da Vinci knew all too well. Yeah, something is going on indeed, but not what you might think. In fact, a mirror has little to do with it. I blame the opaqueness of the material we usually use to write letters on for our lack of immediate insight into the matter.

    Have a look at the drawing in Figure 1. We have replaced a more or less symmetric human by an asymmetric object, resembling the Greek uppercase letter gamma ($\Gamma$) which makes a reflection a bit easier to grasp.

    Figure 1. An asymmetric object standing in front of a mirror. Note that point A and B are not reversed in the mirror image.

    As you can now immediately see, points A and B in our original object have not been reversed in the mirrored object. If you imagine yourself standing behind the original object, looking in the direction of the mirror, point A would be on your left and point B would be on your right. The same is true for the mirrored object: its point A is also on your left and point B is also still on your right. As you have now come to appreciate, this is because the object is reflectedโ€”hallmark of a fine mirror.

    Now imagine the shape of the object being ‘glued’ onto a large sheet as is drawn in Figure 2, representing a large letter printed on a piece of paper.

    Figure 2. Our letter object is glued onto a large piece of paper, representing a printed letter

    So, now we have written a letter on a piece of paper, as it were. The problem is, we don’t see anything in the mirror but the large sheet. What should we do about it? Turn the paper around, you say? Turn around? Okay, cool, sure thing. So, as shown in Figure 3, we rotated the paper. Now, look at what happened to point A and point B. Indeed, A and B have reversed position! And you know why? Because you rotated it about the vertical axis!

    Figure 3. The sheet has been rotated about the vertical axis. The letter is now on the other side but is, for educational purposes, still mildly visible. The positions of points A and B have now reversed.

    How about the mirror? Well, look at Figure 4. It is showing exactly what you are presenting it: the mirror image of a rotated letter. Point A is now on the right, point B is now on the left.

    Figure 4. The mirror is showing you exactly what you are presenting it: a rotated letter

    Conclusion

    And so, letters look weird in mirrors because you rotated them towards the mirror. It is this rotation that makes them look weird.

    Since one usually rotates writings around the vertical axis, it seems like letters, words, and sentences only get reflected in the left-right-direction. They do not the moment they get rotated about another axis.

    A cartoon. The intro states: "Once upon a time in the mirror universe where grown-ups are less informed about how the world works". A child is standing on his head in front of a mirror. He exclaims he's looking so weird right now. Without looking, his parent responds with the false statement that that's what mirrors do: they turn everything around. Meanwhile, the parent is watching a YouTube clip on their phone which seems to be a documentary about a phenomenon already know to ancients for centuries which leaves scientists baffled, struggling for an explanation. The TV shows a documentary about the cutting edge of complementary and integrative medicine based on the well-established principles of quantum mechanics and moves on to quote Albert Einstein (the text stops at this point). At the bottom, a text is displayed in mirror writing, explaining that mirrors do not turn everything around. Mirrors reflect, they don't rotate. They reflect everything you present them. Reflection is not rotation. The text ends with stating that the text itself had been rotated about the vertical axes, hence the mirror writing.

    Two interesting afterthoughts to reflect on. [1] The way you see yourself in the mirror (reflection) is not the way other people see you (rotation). [2] To have letters look as weird as they seem to do in the mirror, you don’t actually need a mirror; just write on a good old transparency, rotate it about the vertical axis, and then hold it in front of you.

    And so, you see, there is no mirror, for it is not the mirror that is reversing the letters: it is only yourself.

    For the sake of completeness, we mention that while mirrors do not reverse left and right, nor up and down, they do reverse front and backโ€”the third of three options in 3D space. But, as letters are symmetric in this direction (they can be regarded as flat, for that matter), this does not usually explain why they look weird, which is why we did not discuss this property. Instead, we emphasised the distinction of a mirror’s reflection (reversal of front and back) from an object’s rotation.

  • Why your coffee does not have tides

    Why your coffee does not have tides


    The Moon orbits the earth and its gravity is causing the tides. But why don’t swimming pools have tides? Or a cup of coffee? Human bodies consist of water, mostly. Aren’t they tidally influenced by the Moon? If you’re asking all these beautiful questions, then what you thought is causing the tides is probably wrong, and here’s why.


    Remember, back in high school, when the science or physics teacher had all the air sucked out of a large, transparent tube which contained a feather and a little steel ball or something like that? And that she asked you to predict which would drop to the bottom first if she would turn the tube upside down?

    Of course, both objects turned out to fall to the bottom at the exact same speed. We learnt it did not matter if the steel ball had more mass than the feather. Earth’s gravity works the same on both. In fact, anything which is being ‘pulled down’ by our planet’s gravity gets to be pulled down at the same rate, no matter how much mass these things have (provided we ignore any form of friction).

    Lunar gravity

    Even though the Moon’s gravity is smaller than Earth’s, the principle is the same. Irrespective of an object’s mass, it falls straight to the lunar surface at precisely the same rate as any other thing. On 2 August 1971, NASA Commander David Scott demonstrated that a feather and a hammer hit the Moon’s soil simultaneously.

    Photo: NASA

    The Moon’s gravity is strong enough to have a noticeable effect on Earth, as we all know. Indeed, it is the reason why our oceans have tides. However, if gravity, whether on our planet or on the Moon, acts the same way on every object irrespective of their mass, how come our bathtub does not experience tides, for instance? Yes, it has less mass, but by Cmdr David Scott’s experiment, that shouldn’t matter. And if the Moon’s gravity is capable of pulling on vast bodies of water such as oceans causing them to rise literally meters high, why does our rubber duck not start levitating up in the air as soon as the Moon rushes past our homes?

    The answer sounds both obvious and contradictory: because the force of the Moon’s gravity is negligibly small, except when it is not.

    The wrong picture

    Let’s have a look at the simplified drawing of Figure 1. Just to make things a little less complicated, we imagine our planet to be covered by water entirely. There are no continents for now.

    We see a schematic drawing of earth and the moon. Earth is covered with water with bulges left and right, representing the two high tides. Point A is the point closest to the moon on the right, located on the surface of the earth in the middle of the bulge on the right. Point B is at exactly the opposite location on the far side of the earth, the most distant point from the moon.
    Figure 1. Earth’s tides and the Moon. (Not to scale!)

    First misconception. Even though, intuitively, it may seem to be the case, the bulge at point A is not because the Moon’s gravity is tugging at it, contrary to popular belief.

    And in many texts, you might encounter the following incorrect explanation for the bulge at point B. ‘The Moon’s pull is smaller at point B than at point A, so, point B stays more or less where it is, while point A gets pulled more towards the Moon. Everything in between A and B gets stretched like chewing gum. So, from the perspective of someone standing (on land) at point B, the water rises there as well.’

    This is also mostly incorrect. It is true, the Moon’s gravitational pull is smaller at B than it is at A. But that is not what is causing the bulge at point B. Not in the direct way as stated here, that is.

    Many a little makes a mickle

    Why don’t we have a look at points C and D in two different, little patches of water in Figure 2? The Moon’s force of gravity acts on these points at a certain angle as is represented by the blue arrows, or vectors. At the same time, the entire earth experiences a slight force towards the Moon as is modelled by the red vector.

    Same schematic as the previous one, but more points are added. Point C is located more or less on top of the earth, a little to the right of the North Pole. Point D is located between the North Pole and point B. Little blue arrows, called vectors, are drawn from points C and D, pointing towards the centre of the moon. A little red vector is drawn at the centre of the earth, pointing to the centre of the moon. These vectors represent the forces acted on these points caused by the moon's gravity.
    Figure 2. The force of the Moon’s gravity acting on points C and D and the entire earth

    So, point C and D undergo two simultaneous forces as is explicitly shown in Figure 3. Note that the blue and red vectors have different directions. Our high school physics or maths teacher then taught us that two or more forces acting on the same point can be modelled as one resultant force.

    Now we need to take two important steps: 1. Newton taught us that a force is an acceleration, so, from now on we will regard the arrows in Figure 3 as being accelerations. 2. To determine the acceleration of the patches of water at points C and D relative to Earth’s surface, we subtract the red vector from the blue vector. What’s left is the green vector, the resultant.

    We see a close-up of points C and D. The red vector, representing the force of the moon exerted on the earth, originates here from point C. The blue vector, representing the force of the moon exerted on point C, still originates from point C. So, we have two vectors coming from point C. A little green vector is drawn between the heads of the other vectors, pointing down, towards the location of where the bulge closest to the moon will emerge. And so, the combination of the two real forces exerted by the moon results in a net force. The exact same procedure has been applied to point D. Only here, the green net force vector is pointed the other way, towards where the other bulge, on the other side of the planet, will emerge.
    Figure 3. The resultant forces are represented by the green vector

    In Figure 3, it is shown how the combination of the two gross forces blue and red yield a net force as represented by the green vectors. Do note, the net forces are what is called apparent forces. Think of a car suddenly accelerating. Relative to the ground, your head is standing still for an instant of time. However, from within the car, your head seems like it is being pushed back by some invisible force. Tides are thus being caused by so-called tidal forces, which are apparent forces.

    So, if we do the same for many other points, you get many green vectors as they are shown in Figure 4. And guess what, all the green (now black) arrows point in a way that look a lot like bulges in the water.

    We see the earth where all the net force vectors are lined up in a way that, together, result in a picture exactly the same as our tides: two bulges on either side.
    Figure 4. An array of net forces (the green arrows are here the black arrows)

    This shows that what actually happens is that every minuscule patch of water gets influenced by a tiny bit of net force in the direction of the places where the bulges will emerge, pushing every other patch in front of it towards the bulges, thereby creating the bulges in the first place.

    Now, in the drawing, all arrows are relatively massive, so we can actually see them. In reality, however, the net forces are tiny. Microscopically tiny.

    And this is the key to solving the paradox. Even though a net force, resulting from the Moon’s gravitational influences, is utterly insignificant on a single patch of the water, the amount of ocean on Earth is quite the opposite of negligible, rendering the sum of all net forces on every cubic patch within the oceanic liquid highly significant, and in some cases, depending on the shape of the land, dangerously significant.

    Conclusion

    The Moon’s influence on tiny things is tiny. It does not noticeably influence your cup of coffee, your body, your bathtub, ponds, and lakes. Any tidal height difference in a cup of coffee could be thinner than a bacterium, the significance of which is immediately squashed by the mere presence of, well, a bacterium in your coffee, practising its back crawl. If your coffee starts to display any tidal effects, prepare for the Apocalypse and/or escaped dinosaurs. Either case, something is really wrong then.

    Even a lake the size of Lake Michigan will only rise a couple of centimetresโ€”easily negated by its murmuring surface on a sunny day in May. So, you can imagine, your body does not feel a thing. The pressure needed for delivering oxygen to your brains alone squashes out every single tidal influence by the Moon, which would have been smaller than a hair’s thickness anyway. If you feel less capable of rational thought, you now know it’s not the Moon. But do check your blood pressure.

    However, in the case of an ocean, a body with many, many, many tiny, watery parts which can roll, slip and slide freely on top of one another, there will be bulges about where the Moon whizzes. However, the swelling occurs by virtue of pushing not pulling, directly.

    In short, a quindecillion minuscule little net forces on every cubic piece of the ocean cause an upward push so the two bulges emerge. Lakes, ponds, bathtubs, human bodies, and coffee mugs do not come close to even a little bit of that amount.

    EDIT: the original article omitted to mention how tidal force is an apparent force, resulting in a fundamental misinterpretation of Figures 3 and 4. This has been corrected.

    Figure 4 is an adapted version (cropped) of the original made by Krishnavedala under CC BY-SA 3.0.

    We did not consider the rotation of the earth, the Coriolis effect, the presence of the sun, the presence of land, etc., just to keep it simple. This changes the situation somewhat, but does not change the gist of it all.

  • Just a minute: Minus minus and negative times negative

    Just a minute: Minus minus and negative times negative


    Minus minus is plus. And negative times negative is positive. Two negatives make a positive. You may have heard or uttered these expressions many times. Even though you will know this already, here you will find an algebraic proof, just for your reference. Requirements: simple algebra from the second year in secondary, high or grammar school.

    Download PDF


    Minus minus is plus

    We all should have learnt in high school that subtracting a negative number is the same as adding the positive version of that number. For example:

    \[ 1 – (-2) = 1 + 2 = 3. \]

    In human English language, it should sound something like: ‘One minus minus two equals one plus two equals three’.

    Now, using just variables instead of numbers, we can write this as

    \[ a-(-b) = a+b, \]

    where $a$ and $b$ are any real number.

    Okay, so let’s prove that, shall we? Or shall we…? Well, not  yet. Let’s first pretend the opposite is true. Suppose,

    \[ a-(-b) = a-b. \]

    Subtracting $a$ from both sides, we get

    \[ -(-b) = -b. \]

    Just to add some clarity, I’m going to slap some brackets around $-b$ on the right hand side:

    \[ -(-b) = (-b). \]

    You can see now, we have a contradiction. I mean, just in case it’s not quite clear yet, let’s suppose $(-b) = c$, so, replacing $(-b)$ with $c$, we get

    \[ -c = c \]

    which is, clearly, in this universe, utterly ridiculous. I mean, $-1=1$? I think not. So, our original statement must be true. Chin-chin, pour some glasses.

    Negative times negative is positive

    We also learnt in high school that multiplying a negative number with another negative number equals some positive number. So, we will prove that

    \[ (-a) \times (-b) = a \times b, \]

    where $a$ and $b$ are any real number. (I put the negative numbers $-a$ and $-b$ between brackets for better visibility, not because they represent any extra information or some sort of an afterthought in the literalistic sense, which this sentence totally does.)

    Mathematicians are true masters of multiplying almost anything at almost any time and even manage to get paid for it. They can truly be a productive lot sometimes. So, of course, for their employer’s money’s worth, they will almost never bother to properly write down the ‘$\times$’-sign. Hence, we write the above equation as if we were actually earning an honest living:

    \[ (-a)(-b) = ab. \]

    The following may seem obvious but bear with us. It’s just the first step. Have a look at this tautology:

    \[ (-a)(-b) = (-a)(-b). \]

    Okay, so far, so obvious. Now, let us add a term without disturbing the essence of the expression:

    \[ (-a)(-b) = (-a)(-b) + 0. \]

    Still a true thing, right? Now, what is also true: anything multiplied by zero equals zero. So, let’s rewrite the expression as follows:

    \[ (-a)(-b) = (-a)(-b) + a\times0, \]

    or, to be a little more pedantic about notation, we could write it more compactly as

    \[ (-a)(-b) = (-a)(-b) + a(0). \]

    Now, let us replace the number 0 by variablesโ€”just the variables we are using here, to be exact. So, let’s say…

    \[ (-a)(-b) = (-a)(-b) + a\underbrace{(b-b)}_{\text{=0}}. \]

    Let’s now get rid of the brackets in the last term. We do this by multiplying out the last term after the ‘+’-sign.

    \[ (-a)(-b) = (-a)(-b) + ab + a(-b). \]

    Let’s swap the order of the last two terms. The next step becomes easier to see. So, swapping term two with term three, yields

    \[ (-a)(-b) = (-a)(-b) + a(-b) + ab. \]

    Now we can comfortably look at the first two terms:

    \[ (-a)(-b) = \underbrace{(-a)(-b) + a(-b)}_{\text{look at this comfortably}} +\ ab. \]

    Remember how to factorise? After factorising, something like $pq + pr$ becomes $p(q+r)$, for example. Guess what, we can do the same for the two terms above, but with $(-b)$ instead:

    \[ (-a)(-b) = (-b)\Big((-a) + a\Big) + ab. \]

    Now, look at the term between the large brackets. This is gorgeous, because, indeed, it amounts to 0. And anything multiplied by 0 is 0. So, what is left, is

    \[ (-a)(-b) = + ab. \]

    Here’s how, old friend. Cheerio.

  • Energy is neither fundamental nor conserved

    Energy is neither fundamental nor conserved


    Sometimes you may have heard someone say that, in the end, ‘everything is energy’. ‘Einstein said himself that mass equals energy, we are energy ourselves, light is energy, and everything in this universe is energy.’ Often, it is represented as the fundamental substance everything is made out of. And energy is conserved. Both statements are incorrect.


    Gottfried Wilhelm von Leibniz

    To get straight to the point: energy is a mathematical concept. It is not a substance and it is not a mysterious ‘elusive something’. Nothing is flowing from one object to the other. It is a number, very useful and ingenious to perform calculations with and base predictions on for the state of a system. It is clever mathematical bookkeeping originating from the seventeenth century polymath Gottfried Wilhelm von Leibniz.

    Here, we must differentiate between physical objects (so, not energy) and properties of those physical objects. Think of properties like position, volume, mass, velocity, and energy. These five proporties are numbers. Mathematical quantities. In high school, we were taught to express quantities in numbers of units, which signified physical phenomena of physical objects, such as, respectively, location, size, inertia, motion, andโ€ฆ what energy signifies, you will read after this.

    Let us take a rolling cannonball A as an example. This physical ball has two measurable properties: a mass A and a velocity A. Suppose, there is another rolling cannonball: mass B, velocity B, only in the opposite direction. They will collide. You may assume that their speeds and the direction of their speeds will have changed after the collision.

    Leibniz noticed that, for each ball its mass multiplied by its squared velocity and then added all together, this total sum before the collision is equal to the total sum after the collision.

    Both the product and the sum are nothing more than a number. The mathematical result of the product of mass and velocity (squared), we call energy. Leibniz, however, did not, but used, rather poetically, the Latin term vis viva, โ€˜living forceโ€™.

    Many years of refinement and extension of the mathematical concept followed. Leibniz’s formulation appeared to be missing a factor of one half, an extension of the vis viva-concept to heat was necessary, and it experienced heavy competition from the conservation of momentum from rival Newton.

    Ultimately, at the beginning of the 19th century, the polymath Thomas Young became the first to use the term energy in written form in his book A Course of Lectures on Natural Philosophy and the Mechanical Arts: In Two Volumes —even though it would still undergo several evolutions. In the end, the concept was not only useful in mechanical and thermal calculations, but also in electrical, magnetic, chemical, and nuclear interactions, for instance.

    Conservation of energy

    Emmy Noether

    Emmy Noether, a mathematical genius, laid the mathematical foundation for the conservation law of energy, among others, which had been formulated a couple of decades before. Thanks to Noether’s theorem, we know why energy is conserved in an isolated system: the laws of nature are so-called time invariant. In other words, for a law of nature it does not matter if it is applied at ten o’clock in the morning or two hours earlier. Whether at four o’clock at night or fifteen minutes later, cannonballs will not collide any differently. Their operations are invariant. If a time-translated, isolated system, such as our cannonballs, works in the exact same fashion as it did before the time translation, we say we have a symmetric situation. And if laws of nature are time symmetric, we can, thanks to Noether’s theorem, derive the law of conservation of energy mathematically.

    Einstein and expansion

    Something many people unfortunately do not know, is that, since Einstein’s general relativity—a little over a hundred years ago, practically at the same time as Noether’s proof of her theorem—the law of conservation of energy does not apply to the observable universe we inhabit, after all.

    That is to say, the law operates just fine at the scale on which we, humans, live our lives on a daily basis. The high school exams are still valid. Architects and engineers can still rely on it. However, at the scale of the observable universe, the one at which cosmologists work, the law does not hold. Spacetime itself is dynamical: it changes over time. Moreover, in 1998, Nobel prize-winning research showed that the observable universe is expanding exponentially. This, too, demonstrates that space itself is not symmetrical over the passage of time.

    The law is thus untenable for the whole observable universe. However, when taken a piece of space and a piece of time small enough, the law works just fine. At this smaller scale, systems appear to be near-isolated from the rest of the universe. Noether’s theorem applies here, and, thus, the law of conservation of energy, which rests on her theorem.

    Not fundamental, but important

    For two reasons, energy cannot be fundamental in a theory of our universe: the concept is a mathematical tool to quantify measurable properties such as mass and velocity and its law of conservation rests on another theorem, while, at the same time, it has been proven not to be conserved, about a hundred years ago.

    Even though not an invisible, flowing substance or some other mysterious fundamental quantity, it is, nevertheless, highly useful in diverging areas such as fluid dynamics, statistical mechanics, astrophysics, nuclear physics, and quantum physics, even just to simply replace an intricate formulation such as

    \[ \frac{mc^2}{\sqrt{1-\dfrac{v^2}{c^2}}}, \]

    by

    $E$

    for energy. Eh, ‘energy’.

    Photo by ESA/Hubble/NASA. A Hubble Space Telescope image of Galaxy cluster Abell 2537. The amount of gravity, that is, warping of spacetime, caused by this galaxy is visible through the bending of the light of stars and galaxies behind Abell. The galaxy works as a lens. All is predicted by Einstein’s General Relativity.

  • When and why do you multiply probabilities?

    When and why do you multiply probabilities?


    At high school you may have been taught that, sometimes, you have to multiply probabilities. We briefly discuss when and why you do this.

    Download PDF


    First a few notes on the notation of probabilities. When throwing with a dice, the event of throwing a six is 1 of 6 possibilities. We write this as a fraction, 1/6, or

    \[
    \frac{1}{6}.
    \]

    We then say there is a probability of 1 out of 6 to throw, for instance, a 6. The probability is 1/6, one sixth.

    This also means that the probability of throwing a number—this can thus be any number: 1, 2, 3, 4, 5 or 6—is equal to \[ \frac{6}{6} = 1. \]

    If you throw a dice, the probability is 1 for throwing a number, or 100%. In other words, if something is 100% certain to happen, the probability is 1. And if something is less certain to occur, less than 100%, the probability is an n’th part of 1.

    Lastly, an important announcement on multiplying by a fraction: if you calculate an n’th part of something, for instance, 16, you can write this in two ways. You divide 16 by 2 or you multiply 16 by 1/2. It is the same. That is: \[ \frac{16}{2} = 16\times\frac{1}{2} = 8. \]

    Two coins

    Suppose, you throw euro #1 into the air. It is going to be either heads or tails. In Figure fig:figure1, this is represented schematically. The probability of throwing heads is 1/2. The odds of throwing tails is 1/2.

    Figure 1

    Imagine throwing euro #1 and euro #2 into the air. This is represented in Figure 1.

    Now, ask yourself the question: of all the times I threw heads with euro #1, how many times would I have thrown euro #2? The answer is that half of the time euro #1 became heads and half of that time europ #2 became heads.

    What is half of a half? This is
    \[
    \underbrace{\frac{1/2}{2}}_\text{half of a half} = \frac{1}{2} \times \frac{1}{2} = \frac{1}{4}.
    \]

    Figure 2

    Part of a part

    See Figure fig:figure3. Suppose, we throw two coins 16 times. Suppose, coin number 1 turns out heads half the time; we signify this with blue circles. The question is how many times that coin number 1 is heads, do we throw heads with the second coin? This is, again, half. Half of half, that is. We paint this green.

    Of the total amount of throws, what part is green? Half (4) of half (8) of the total (16), so 4 out of 16, or 1 out of 4. So, what is the probability of throwing green (coin number 2 is heads) if you throw blue (coin number 1 is heads) half of the time. \[ \frac{1/2}{2} = \frac{1}{2}\times\frac{1}{2} = \frac{1}{4}. \]

    Figure 3

    A euro and a dice

    Another example. Suppose, you throw a euro and a dice into the air. The probability distribution of heads and tails is 1/2, as we know. In the case of the dice this is different: it can turn out to be 1, 2, 3, 4, 5 or 6. So, the probability of throwing a six is 1/6.

    Of all the trials where the euro turned out to be heads—which is half of the total amount of trials—how many times would you have thrown a 6 with the dice? That is, thus, 1/6th of one half of the total amount of trials. Or, \[ \frac{1/2}{6} = \frac{1}{2} \times \frac{1}{6} = \frac{1}{12}. \]

    Conclusion

    Suppose, that the probability is 2/3 for event $A$ to happen, the probability is 1/6 for event $B$ to occur, and 4/5 that even $C$ will happen, then the probability of the combination of the events $A$, $B$, and $C$ to occur is equal to \[ \underbrace{\frac{2}{3}}_A \times \underbrace{\frac{1}{6}}_B \times \underbrace{\frac{4}{5}}_C = \frac{8}{90} = \frac{4}{45} \approx 0.088\dots, \] which is 8.9% rounded to one decimal.

    To know what the probability of a combination of events occurring is, we calculate the n’th time of an n’th time. And the n’th time of an n’th time (of an n’th time, etc…) is the same als multiplying the two (or more) fractions.

  • The riddle of birthdays

    The riddle of birthdays


    Probabilities can be hard to grasp. For instance, what are the chances that among a birthday party’s attendants two or more people will have their birthdays on the same day? Probably better than you might expect.

    Download PDF


    Since the day she was born, every year, my mother’s birthday has been on 1 January. This year, she celebrated her twelfth jubilee year in a cosy party room filled with about fifty people.

    Being the life and soul of any party, during my little talk, I presented the guests the fact that the probability of my mother’s day of birth being 1 January equalled 1/365. As most years consist of 365 days, I left leap years out of consideration. I also assumed a uniform distribution of birthdays in a year as this makes it easier to perform further calculations.

    Then I asked what the probability was for my father to be born on 8 July, given that there 365 days to choose from. The answer was, again, 1 out of 365, or 1/365. Of course, in this respect, a particular day is not more special than another particular day other than the cultural significance we assign to some.

    Then I asked the crucial question: what is the probability that two or more people in this room share the same birthday? Of course, irrespective of their year of birth. It was purely about the day of the year.

    In other words, there are 365 days in a year and we have 50 people whose birthdays have spread over those 365 days. What is the probability that two (or more) birthdays fall on the same day?

    Here, I am increasing the party fun by putting forward a maths riddle during my little talk.
    Here, I am increasing the party fun by putting forward a maths riddle during my little talk.

    Sometimes, people think of an example with dice. Suppose, you have two dice. The probability of throwing a six is 1/6, which is the same for throwing a six with the other dice. The chances of throwing sixes with both dice is, thus, 1/6 $\times$ 1/6 = 1/36. Logically, the probability is smaller than throwing a six with one dice. (Read When and why do you multiply probabilities?) Many people argue that the probability of two people having their birthday on 8 July, for instance, is therefore equal to 1/365 $\times$ 1/365 = 1/133225; which is, therefore, a very small probability. This would be in accordance with many people’s intuition: it would be highly unlikely if two people, within a group of fifty people, would share their birthday, wouldn’t it?

    Others think of 50 marbles in a jar with 365 marbles. You draw one marble out of the jar en put it back again. You then shake the jar. Again, you draw a marble out of the jar. What is the probability you draw the same marble out of the jar? This way, people get the answer of 50/365.

    But no, both strategies are incorrect. In reality, the probability is 97%, rounded to the nearest integer percentage. Therefore, I would want to bet a good bottle of wine on this.

    The calculation

    Often, in mathematics, it is easier to explore the opposite situation. Let us proceed accordingly. The reverse situation is that no one shares their birthday. Let us look at this more closely. What is the probability no one shares their birthday?

    We have 365 days. We have 50 humans. What is the probability that human number 1 in the group has their birthday on a day? Mind the phrasing of the question. A day, not a specific day. Hence, we are not asking what the probability is of being born on, for instance, 8 July. We are asking ourselves what the probability is of human number 1 being born on one of those 365 days. Well, that is 365 out of 365 days, or 365/365, or $365:365=1$, or 100%.

    Now, what is the probability that human number 2 in the group has their birthday on a day, but not the same day as human number 1 has theirs? Therefore, the possibilities for human number 2 to have their birthday are one fewer than 365, which is, perhaps not surprisingly, 364. Otherwise, both human number 1 and 2 could have had their birthdays on the same day. So, the probability that human number 2 is having their birthday on another day is 364 out of 365, or 364/365, or $364:365=0.997\dots$, or 99.7%.

    And so, what is the probability that both have their birthday on different days? This is

    \[ \frac{365}{365} \times \frac{364}{365} = 1 \times 0.997\dots = 0.997\dots, \]

    or 100% rounded to the nearest integer percentage. In other words, de chances are practically non-existent for them having their birthday on the same day.

    Yet, what happens if we involve a third human? What is the probability for human number 3 having their birthday on a different day than those of humans number 1 and 2? This is then 363 out of 365, or 363/365, or $363:365=0,994\dots$, or 99.4%. And so, what are the odds that humans number 1, 2 and 3 all have their birthdays on different days? This is then

    \[ \frac{365}{365} \times \frac{364}{365} \times \frac{363}{365} = \frac{132132}{133225} = 0.991\dots, \]

    or 99%, rounded.

    Just to make sure, let us involve a fourth human. The probability for human number 4 to have their birthday on a different day than those of humans number 1, 2 and 3, is 362/365, or $362:365=0.991\dots$. And so, what is the probability that humans numbers 1, 2, 3 and 4 have their birthdays on different days? This is then

    \[ \frac{365}{365} \times \frac{364}{365} \times \frac{363}{365} \times \frac{362}{365} = \frac{47831784}{48627125} = 0.983\dots, \]

    or 98%, rounded. We can now clearly observe the decreasing probability of humans having their birthdays on different days with each addition of humans.

    Imagine we would continue this process until human number 50. The probability that human number 50 has their birthday on any other day than the rest of the 49 preceding humans, then becomes 316/365. So, the probability that all fifty humans have their birthdays on different days is calculable as follows:

    \[ \frac{365}{365} \times \frac{364}{365} \times \frac{363}{365} \times \frac{362}{365} \times \dotsm \times \frac{317}{365} \times \frac{316}{365} = 0.029\dots, \]

    that is, only 2.9%!

    So, now we have our answer. The opposite situation, i.e. the probability that not all fifty humans have their birthdays on different days, is the reverse of 2.9% and this is 97.1%. Or, 97% rounded.

    Among the fifty guests, no fewer than six people turned out to share their birthday. In other words, we found three ‘pairs’ sharing their birthday.

    By the way, among a group of 23 people, the probability is already 0.504 (so, just over 50%) for two (or more) people to share their birthday. The odds grow favourably quickly.

    Might you want to read more, this phenomenon rests on the pigeonhole principle or Dirichlet’s box principle—this nineteenth century German mathematician was probably the first one to formalise it. Happy Googling! (Though we highly recommend DuckDuckGo.com.)

    You can use our calculator to quickly calculate the probability for a number of people you specify.

    Photo of the birthday cake by Will Clayton under CC BY 2.0.

  • Happy New Year: Earth is Amazing

    Happy New Year: Earth is Amazing


    NASA recently published a 4K version of an audiovisual amalgam of the iconic Earthrise photo made by Apollo 8’s astronaut William Sanders on the 24th of December 1968, a 3D mapping of lunar photography, and the voice-recording of the crew onboard the spacecraft.


    It allows us to witness in real time the events unfolding leading up to the moment it was captured. We can hear Anders’ crew-mate Borman joking: ‘Hey, don’t take that, it’s not scheduled.’ We are glad Anders was not the type to take everything literally.

    In an interview with The Guardian, Anders noted his experience had even changed his religious views, in fact, undercutting them.

    Well, while we are not concerned with whether it changes one’s existential views on life and the universe or not, we do hope you enjoy witnessing their voice-recorded awe and amazement for this precious little blue ball and its smaller grey companion whizzing around the big white ball for a whole new year.

    We wish you a successful, loving, and all-along-the-line gorgeous new trip in the vast emptiness of the fabric of the cosmos.

    Featured image: NASAApollo 8 Crew, Bill Anders; Processing and License: Jim Weigang

  • Boxing day: Marie and Pierre Curie announce the discovery of radium

    Boxing day: Marie and Pierre Curie announce the discovery of radium


    Today, that is, on the 26th of December, Marie Curie, her husband Pierre, and Gustavรฉ Bemont announced in the Comptes rendus de l’Acadรฉmie des Sciences in 1898 that they discovered a new element which they proposed to name radium.


    Nowadays, an isotope of radium is of medical use against metastatic bone cancer, relying on the isotope’s release of helium nuclei (so-called ฮฑ-decay) which destroy cells within a relatively small volume of tissue.

    An English translation of the original article is available online. The original can be read here (pp. 1215-1217):

    Featured image: Marie and Pierre Curie, and Gustavรฉ Bemont in their laboratory in Paris, 1901. Welcome Images, CC-BY 4.0

  • The Arts in Physics: a short film of freezing soap bubbles

    The Arts in Physics: a short film of freezing soap bubbles


    It’s mesmerising. As soap bubbles freezeโ€”an incredibly delicate processโ€”the camera of Don Komarechka recorded each little detail beautifully. According to the video description, it took about 400 attempts to produce the short film you’re seeing here. Originally licensed to the BBC for their Forces of Nature documentary series, two years after the fact, we are happy to see it out in the open. Physics and art brought together wonderfully. Merry Christmas!

  • What is a spacetime interval?

    What is a spacetime interval?


    Einstein and collaborators taught us that space and time are not fixed quantities. They can stretch and contract. They vary. There is one thing, though, that does not vary. It is the invariance of the spacetime interval.

    Download PDF


    Spatial interval

    Suppose, a photon is emitted from origin $O$ and travels to point $F$ as depicted in Figure 1. Let us write down the expression for its distance-squared, $d(O,F)^2$, in terms of the other distances using the good old Pythagorean theorem:

    $ d(O,F)^2 = d(O,A)^2 + d(A,B)^2 + d(B,F)^2. $

    Figure 1 A photon travels from O to F in a three-dimensional space

    We can also write the previous expression in terms of their distance from $O$. We then write the following:

    \begin{align}
    F &= (x, y, z),\quad O = (0,0,0), \\
    d(O,F)^2 &= (x-0)^2 + (y-0)^2 + (z-0)^2, \\
    \therefore d(O,F)^2 &= (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2. \label{eq:distance O-F}
    \end{align}

    As the axes of the space in Figure 1 are spatial and Euclidean, $d(O,F)^2$ is called a spatial or Euclidean interval. It can also be thought of as a rectangular cuboid represented by its space diagonal $d(O,F)$, tracing out a region of 3D space.

    In the real world, to make sure we meet at the correct place, we could, for instance, give the following coordinates: 1 Einstein Drive, 2nd floor. Think of Einstein Drive as some place along the the $x$-axis (next to $x$-axis places like Battle Road, Mercer Road), number 1 as some place along the $y$-axis, and 2nd floor as some place along the $z$-axis.

    What we still need, though, is an extra bit of information: when do we meet?

    Time interval

    Suppose, Figure 2 shows a timed series of our photon on its way to point $F$ and beyond. It demonstrates that we live in a world where we do not just need three spatial coordinates, but also a time coordinate. It is only logical to not just tell the people you are supposed to meet, where in space you will be, but also when in time you will be there.

    Our photon $P$ flies through $F$ at $t=3$. This entails that the time coordinate of the event that the photon reaches $F$ is
    $ t_{F}=3. $

    Assuming that at $t=0$, photon $P$ is at the origin,

    $ t_{O}=0, $
    then we can write for the temporal interval between the photon leaving $O$ and reaching $F$:

    $ \Delta t_{OF} = t_{F} – t_{O} = 3 – 0 = 3. $

    Figure 2 A photon travels from O to F in a three-dimensional space over a period of time

    Time to distance unit conversion

    As the previous two sections showed, we need four coordinates to describe an event, for instance, the event where photon $P$ reaches $F$. The three spatial distances are measured in a unit of distance, usually, metres. The one temporal distance is not a distance in the traditional sense and is usually expressed in seconds. This makes it difficult to make sensible comparisons.

    To convert the time-units to distance-units, we multiply by a constant of nature, the speed of light $c$, which, by Einstein’s second postulate [1], happens to be invariant: no matter which frame of reference you choose, the speed of light is constant. For a longer description of this conversion, read section 4.3 of Deriving the Lorentz transformations from a rotation of frames of reference about their origin with real time Wick-rotated to imaginary time. We conclude that our time interval becomes a temporal distance:

    $ \Delta t \mapsto c\Delta t. $

    Spacetime interval

    In Figure 3, we left out the spatial $z$-axis and replaced it with the temporal $ct$-axis (which is thus time expressed in distance-units) in order to make a comprehensible drawing on a flat surface. In reality, of course, the photon still moves in the $z$-direction as well. (We have thus not yet been successful to draw a four-dimensional object on a flat surface.) Mind the unit vector diagram top right and the points of distances $\Delta x$, $\Delta y$, and $c\Delta t$. Then think, really hard, of an added fourth spatial distance $\Delta z$, somewhere.

    To calculate the spatial distance $d(O,F)$ for our photon, we look again at Equation \eqref{eq:distance O-F}:

    $ d(O,F)^2 = (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2.\quad\eqref{eq:distance O-F} $

    Since we know that speed, in general, is calculated through $v = \Delta x / \Delta t$, where $x$ is the travelled distance in one direction, and $v = c$ for our photon, we can write for the travelled distance of our photon from $O$ to $F$:

    \begin{align} d(O,F) &= v\Delta t, \\ d(O,F)^2 &= (v\Delta t)^2, \\ \therefore d(O,F)^2 &= (c\Delta t)^2. \end{align}

    This is becoming interesting, since $(c\Delta t)^2$ is also (the square of) the temporal distance. If we substitute Equation \eqref{eq:distance O-F} into this last equation, we get

    $ (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2 = (c\Delta t)^2. $

    If we rearrange this a little bit, we get

    \begin{equation}
    – (c\Delta t)^2 + (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2 = 0. \label{eq:spacetime-homogeneity}
    \end{equation}

    While this may seem nice and simple, the question we should be asking is, what is zero? If we know the answer to that, we know the answer to what all the terms are on the left-hand side of the equals sign.

    In physics and mathematics, whenever something equals zero, something special is going on: it may entail a certain system in a certain configuration that is stable, static even, it may point to constant motion, an energy well, an attractor, a root, a conservation law, homogeneity, or a minimum or maximum of some kind.

    In general, it means that there is a certain kind of symmetry at play, which in turn means that something is conserved. There are beautiful, deep insights to be made as Emmy Noether showed us [2], and her genius deserves nothing less than an entire series of articles on their own.

    However, for now, let us conclude that independent of which coordinate system one uses, rendering different values for $\Delta x$, $\Delta y$, $\Delta z$, and even $\Delta t$, as we have come to learn from the Lorentz transformations, the sum of all these variances remains invariant. The zero points to the fact that irrespective of its four moving parts—no matter what frame of reference you prefer—the resultant is a constant, i.e. invariant.

    The quantity on the left-hand side has a name; it is called the spacetime interval and is denoted by $(\Delta s)^2$. The $s$ stands for ‘separation’. It is about the separation between events. If we had used the word distance, it might have had inadvertently referred too much to a spatial distance, hence, we use separation, $(\Delta s)^2$. And so, the spacetime interval is usually written:

    \begin{equation}
    (\Delta s)^2 = – (c\Delta t)^2 + (\Delta x)^2 + (\Delta y)^2 + (\Delta z)^2.
    \end{equation}

    The signs before the terms may have been flipped in some texts, but important to note is that, while time has been made comparable to space unit-wise by multiplication by $c$, you can still see that time has a special place in the interval of the fabric of the cosmos.

    Figure 3 A spacetime diagram with two spatial dimensions and one temporal dimension.

    Featured image: Klaus P. Rausch

    References
    1. A. Einstein, Zur Elektrodynamik bewegter Kรถrper, Annalen der Physik 322(1905), no. 10, 891—921.
    2. E. Noether, Invariante Variationsprobleme, Nachrichten von der Gesellschaft der Wissenschaften zu Gรถttingen, Mathematisch-Physikalische Klasse 1918(1918), 235—257.
  • Deriving the Lorentz transformations from a rotation of frames of reference about their origin with real time Wick-rotated to imaginary time

    Deriving the Lorentz transformations from a rotation of frames of reference about their origin with real time Wick-rotated to imaginary time


    Well-known for their central role in Einstein’s Special Relativity, the Lorentz transformations are derived from the rotation of two frames of reference in standard configuration while time is taken to be an imaginary unit of spacetime. This is rarely seen in the wild. Not many undergraduate textbooks or online texts show the details of the working. Hence, this article.

    Download PDF

    Introduction

    One might think this means that imaginary numbers are just a mathematical game having nothing to do with the real world. (โ€ฆ) It turns out that a mathematical model involving imaginary time predicts not only effects we have already observed but also effects we have not been able to measure yet nevertheless believe in for other reasons. So what is real and what is imaginary? Is the distinction just in our minds?

    S. Hawking[1]

    Even though there are many derivations of the Lorentz transformations to be found in textbooks, in syllabi, and online, to me, one of the most elegant remains the version Henri Poincarรฉ once alluded to[2], which Hermann Minkowski then toyed with a bit further—to put it unreasonably mildly—in what we now call Minkowski space, but is rarely expounded in the aforementioned places.

    Henri Poincarรฉ noted that, when the time axis of the two coordinate systems has been made imaginary, i.e. the imaginary axis in the complex plane, the transformations set forth by Hendrik Lorentz pop out automatically after a rotation of two reference frames in that complex plane.

    In this document, we show how this is done. We assume the reader is familiar with complex numbers.

    The aim is to derive the following set of Lorentz transformations:

    \begin{align}
    t’ &= \frac{t-vx/c^2}{\sqrt{1-v^2/c^2}}, \label{eq:Lorentz t-prime} \\
    x’ &= \frac{x-vt}{\sqrt{1-v^2/c^2}}, \label{eq:Lorentz x-prime} \\
    y’ &= y, \\
    z’ &= z,
    \end{align}

    where $(t,x,y,z)$ and $(t’,x’,y’,z’)$ are the coordinates of an event in two frames. The primed ($’$) frame is, seen from the unprimed frame, moving with speed $v$ in the $x$-direction. The speed of light in a vacuum is denoted by $c$. As a side note, the recurring term $(\sqrt{1-v^2/c^2})^{-1}$ is called the Lorentz factor and it is usually denoted by the letter $\gamma$.

    Standard configuration

    Figure 1: Two frames of reference in standard configuration. A two-dimensional depiction of Hermann Minkowski’s frame M and Albert Einstein’s frame E in standard configuration: the latter moves at speed v relative to the first in the direction of x. There is no motion in either the y- or z-direction. Note that some time has passed in this diagram. At time t=0, however, their origins were equal. In other words, at temporal coordinates t=t’=0, their spatial coordinates were the same, thus x=x’=0.

    Suppose, Hermann is standing still on the ground. Albert is driving his car and moves away from Hermann at speed $v$. We then have two frames of reference. There is Hermann’s frame $\mathcal{M}$ (the ground), with its origin $O$ at Hermann’s feet on the ground. And there is Albert’s frame $\mathcal{E}$ (the car), with its origin at Albert’s bottom on his chair. Their frames of reference are said to be in standard configuration as depicted by Figure 1. This means that at time $t=0$ in frame $\mathcal{M}$ where $x=0$ as well, the time $t’=0$ and position $x’=0$ in frame $\mathcal{E}$, too, and that one frame is in uniform (constant) motion relative to the other. In other words, $\mathcal{M}$ and $\mathcal{E}$ are said to be synchronised when the spacetime coordinates

    \[ (t,x) = (t’,x’) = (0,0). \]

    Of course, in the real world, there are four spacetime coordinates for each frame, i.e. $(t,x,y,z)$ and $(t’,x’,y’,z’)$, but to make our calculations a little bit easier, we consider the temporal coordinate $t$ (and $t’$) and spatial coordinate $x$ (and $x’$) only.

    So, looking at Figure 1, we can see from Hermann’s point of view โ€“ standing in the origin $O$ of $\mathcal{M}$ โ€“ that $\mathcal{E}$’s origin $O$ moves at speed $v$ relative to the $x$-axis of $\mathcal{M}$. Speed $v$, of course, just means that $\mathcal{E}$ moves at a certain amount of units of $x$ (say, metres) per a certain amount of units of $t$ (say, seconds). This is nothing new, but it is for our derivation of the Lorentz transformations important to repeat our secondary education for a little bit:

    \[ v = \frac{\Delta x}{\Delta t}, \]

    in Hermann’s frame of reference $\mathcal{M}$. More or less conversely, if we want to calculate how many spatial units frame $\mathcal{E}$’s origin has moved from frame $\mathcal{M}$’s origin, we rewrite the last equation into the perhaps more familiar law of uniform motion:

    \begin{equation}
    \Delta x = v \Delta t,
    \label{eq:x=vt}
    \end{equation}

    in Hermann’s frame of reference $\mathcal{M}$.

    As a side note, do realise that to Albert, his car is not moving at all; he is sitting in it. (Rather, it is the rest of the world that is moving with respect to his car and himself.) If the car were moving with respect to Albert, an accident with potentially serious consequences would be impending. So, for Albert’s sake, his speed within his own frame $\mathcal{E}$ (the car) is expressed as $v’=0$, as long as he stays put and buckled up in his chair. And so, the law of uniform motion of Albert, within his frame $\mathcal{E}$, becomes:

    \[ \Delta x’ = v’ \Delta t’ = 0 \Delta t’=0. \]

    Invariances

    If $\mathcal{E}$ is in constant motion with respect to $\mathcal{M}$, in one direction, the $x$-direction, as expressed in Equation \eqref{eq:x=vt}, then, mathematically, we call this a geometric translation in the $x$-direction. In physics, it is called a translational motion in the $x$-direction.

    Figure 2: Albert fires a photon. At time t=t’=0, Albert fires a photon P into direction x. Both the photon and Einstein’s frame E move into that same x-direction.

    Suppose, at time $t=t’=0$, Albert activates his special on-board photoelectric cannon, firing exactly one photon $P$ in the $x$-direction. Figure 2 shows how the photon is travelling through the spaces of both frames of reference.

    Looking at the diagram, we see that the spatial coordinates in the $y$-direction remain unchanged, $y=y’=0$, so we leave this out of our equations further on, to keep it simple. However, since $\mathcal{E}$ is moving with respect to $\mathcal{M}$ in the $x$-direction, we do know that $x\neq x’$ for $t>0$. And since we do not know for certain that $t=t’$ for $t>0$, only that $t=t’=0$, we will have to conclude that the position of $P$ differs:

    \begin{equation}
    \begin{aligned}
    \text{in Albert’s }\mathcal{E}\text{: }P &= (t’,x’), \\
    \text{in Hermann’s }\mathcal{M}\text{: }P &= (t,x).
    \end{aligned} \label{eq:coordinates of P}
    \end{equation}

    Fortunately, accepting Einstein’s Voraussetzungen[3], we know that the speed of light, $c$, is the same for every frame of reference. Using Equation \eqref{eq:x=vt}, $x=vt$, and the fact that $v=c$, in this case, we can write for the distance travelled of photon $P$ โ€“ the yellow line in the diagram โ€“ in the coordinates of the respective frames of reference:
    \begin{align}
    \text{in Albert’s }\mathcal{E}\text{: }\Delta x’ &= c\Delta t’, \\
    \text{in Hermann’s }\mathcal{M}\text{: }\Delta x &= c\Delta t.
    \end{align}

    As it is possible for any coordinate system to have points which lie on the negative side of the origin of a spatial dimension such as $x$ in our case, and thus for light to travel in the negative $x$-direction, we simply square both equations to always obtain a positive value.

    \begin{align}
    (\Delta x’)^2 &= (c\Delta t’)^2, \\
    (\Delta x)^2 &= (c\Delta t)^2.
    \end{align}

    If we then rearrange this,

    \begin{align}
    (\Delta x’)^2 – (c\Delta t’)^2 &= 0, \\
    (\Delta x)^2 – (c\Delta t)^2 &= 0,
    \end{align}

    we see that both are equal to zero, allowing us to write

    \begin{equation}
    (\Delta x’)^2 – (c\Delta t’)^2 = (\Delta x)^2 – (c\Delta t)^2.
    \label{eq:interval}
    \end{equation}

    This is a beautiful result, because it tells us that no matter what frame of reference you happen to be in, besides $c$, Albert and Hermann agree on the quantity $(\Delta x)^2 – (c\Delta t)^2$, despite the fact that the coordinates of $P$ are not necessarily the same in every frame of reference as we saw in \eqref{eq:coordinates of P}. In other words, both $c$ and $(\Delta x)^2 – (c\Delta t)^2$ are said to be invariant.

    You might wonder, what is this invariant quantity $(\Delta x)^2 – (c\Delta t)^2$, exactly? Well, this will be discussed in another post called What is a spacetime interval? And now, we might have just told you what it is. Anyway, let us move on to deriving the Lorentz transformations, and just keep in mind that $(\Delta x)^2 – (c\Delta t)^2$ is a wonderfully invariant quantity, equal in both frames of reference. Let us move on to imaginary time.

    Wick rotation and imaginary time

    Number sets

    Figure 3: real number line. A segment of the real number line, the set R of all real numbers, which goes on to infinity on either side.

    As many of us should know, Figure 3 depicts (a segment of) the real number line, that is the set $\mathbb{R}$ of all real numbers. It formed the culmination of all the previous extensions of the then-known set of numbers. Starting with the natural numbers, a set usually denoted by $\mathbb{N}$, containing all positive integers, arisen from the natural act of counting, the numeric repertoire was then extended by the notion of negative integers. Instead of just 1,2,3, we could now also count to -1,-2,-3 etc. This extension is denoted by $\mathbb{Z}$. Needless to say that $\mathbb{N}\subset\mathbb{Z}$, but we just did anyway.

    Of course, some people were clever, acknowledging the need for another extension: numbers which represented ratios, better known as rational numbers, such as 1/2, 1/-3, 1/4, -1/100, in other words, quotients of two integers. These numbers would sit in-between the integers in $\mathbb{Z}$. The symbol is $\mathbb{Q}$, and it is superfluous to add that $\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}$.

    While specified on a tablet, found in Susa (Iraq) in 1936, dated as used by Babylonians around 2000 BCE, that

    \[ \frac{3}{\pi}=\frac{57}{60}+\frac{36}{(60)^2}, \therefore \pi = \frac{25}{8}=3.125, \]

    it wasn’t until 1761 that a proof that $\pi$ is irrational was found by Johann Heinrich Lambert[3] meaning that it could not be constructed by any ratio of integers, as were many other numbers, such as $\sqrt{2}$. And so, yet again, an extension of the existing number line was needed. This was the aforementioned line representing the set $\mathbb{R}$, or, to be precise, $\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}$.

    And then, in the 16th century, people such as rivals Tartaglia and Cardano independently recognised that solutions to cubic equations sometimes required the manipulation of square roots of negative numbers, such as $\sqrt{-1}$. Later, Bombelli developed proper operations such as addition and subtraction. A whole slew of subsequent mathematicians then developed over several decennia what is now known as the complex plane or gaussian plane[5], representing the set $\mathbb{C}$, extending the real number line with an imaginary axis with multiples of the imaginary unit $i=\sqrt{-1}$. (It is, obviously, the solution to the quadratic $x^2+1=0$.) We realise mentioning that $\mathbb{N}\subset\mathbb{Z}\subset\mathbb{Q}\subset\mathbb{R}\subset\mathbb{C}$ is utterly redundant at this point.

    Translation and rotation

    Figure 4: Number sets. Every consecutive number set is an extension of the previous one. We can move from a simpler set to a more complex one, for instance, by simply multiplying our current position by a number only present in the more complex set. Although it seems like we are tumbling from one set to another, we are really just ‘sliding left or right’, one-dimensionally, on the number line of the more complex set. This sliding is called a translation. Note: the amount of ticks in Q is much larger, but for obvious reasons of legibility, we only ticked every 1/2-ratio.

    Let us have another look at the natural number line of $\mathbb{N}$. If we would want to convert the number 1 to a number that could not exist in $\mathbb{N}$ but could exist on the integer number line of $\mathbb{Z}$, let us then simply multiply the natural number 1 with a number from $\mathbb{Z}$, the negative integer $-1$. Since $1\times-1=-1$, we have transitioned from $\mathbb{N}$ to $\mathbb{Z}$. We ‘slid’ from 1 to $-1$, albeit in a different number set, which, mathematically, is the same as a translation by $-2$. This is easily expressed as starting from position 1, adding $-2$, and ending up at position $-1$ on the number line of, at least, $\mathbb{Z}$ (but not $\mathbb{N}$): $1+-2=-1$. Figure 4a aims to depict this.

    We can do the same with moving from position $-1$ in $\mathbb{Z}$ to a number not in $\mathbb{N}$, nor in $\mathbb{Z}$, but at least in $\mathbb{Q}$ by simply multiplying by a fraction, such as $-1/2$, which is also a number not in $\mathbb{N}$, nor in $\mathbb{Z}$. This is, again, actually a translation, though now by adding $3/2$: $-1+3/2=1/2$. Figure 4b aims to depict this.

    Similarly, transforming from position 1 in $\mathbb{Q}$ to $\mathbb{R}$, we multiply by, for instance, $\sqrt{2}$, which exists in $\mathbb{R}$ but not in $\mathbb{Q}$, and so, the result, $1\times\sqrt{2}=\sqrt{2}$ is in at least $\mathbb{R}$ but not in $\mathbb{Q}$, nor in $\mathbb{Z}$, nor in $\mathbb{N}$. The result is also a translation of $1+(\sqrt{2}-1)=\sqrt(2)$ in $\mathbb{R}$. Figure 4c aims to depict this.

    Figure 5: Wick rotation of 1 and real time. Compared to the increasing complexity of the number lines in Figure 4, this one is the most complex so far. Two Wick rotations into the complex (number) plane C, where (a) the Re-axis is the real number line of the set R and the Im-axis is the imaginary unit line of the set C. Note that a complex number consists of both: a real part and an imaginary part. For instance, the complex number z is written in the form z = a + bi, with i = โˆš-1. The real part of z is Re(z) = a, and the imaginary part Im(z) = b. And so, z = 1 + i, z = 1/2 + 3i, z = 3, are all complex numbers, where the latter has an imaginary part of Im(z) = 0, which we simply leave out as 0i = 0. A complex number is thus two-dimensional, embedded in a plane with a real axis and an imaginary axis. (b) Wick-rotating all real numbers on the time axis to the imaginary axis, transforms real time into imaginary time.

    Note that, so far, the transformation of 1 or another number has involved a simple ‘sliding’ motion on the number lines, that is, one-dimensionally. Every time a new kind of number was introduced โ€“ the negative integers, ratios, and, lastly, the real numbers โ€“ a new set of numbers was created, and the number line evolved from discrete ($\mathbb{N}$) to a line continuum $\mathbb{R}$. The question is, what would be the next extension and what would it look like?

    As stated earlier, in the 16th century it became clear that a new type of number was necessary to solve a slew of quadratic equations. Owing to people such as Wallis, Wessel, Argand, Buรฉe, Mourey, Warren, Franรงais, Bellavitis, Gauss, and Euler[5][6], the idea to extend the real number line of $\mathbb{R}$ with an imaginary, perpendicular number line came to fruition. This created the so-called complex (geometric) plane, sometimes called the $z$-plane, Gauss plane or Argand plane. It is important to note that transforming a real number in $\mathbb{R}$ to a complex number in $\mathbb{C}$ involves not a translation but a rotation. Multiplying a real number in $\mathbb{R}$, say 1, by a number that only exists in $\mathbb{C}$, say $i$, is the same as geometrically rotating our position 1 on the real axes by $\pi/2$ onto a position $i$ on the imaginary axis as is depicted in Figure 5a.

    What if we did this with the entire real time axis in a space-time diagram as shown in Figure 5b? Every element of real time $t$ is multiplied by $i$. In other words, every part is rotated in the complex plane to become an entire imaginary axis of time $it$. This procedure is called a Wick rotation, named after theoretical physicist Gian Carlo Wick, who described such a procedure to solve problems in quantum and statistical mechanics[7].

    This seems promising and is what Henri Poincarรฉ alluded to fifty years earlier. Before we continue, we have to do one other little thing. It has something to do with units of measurement.

    Minkowski diagrams

    Figure 6: The distance-time-diagram we all grew up with. The independent variable time t as the x-axis, and the dependent variable distance, x, as the y-axis. Four particles are travelling through time and (one-dimensional) space, each with its own distance function of time, that is, each with its own speed. Note that P travels the most amount of distance over the same period, in other words, it is the fastest. Note that S travels exactly zero distance in that same amount of time.

    We all grew up learning to read and use a type of diagram as depicted in Figure 6 during our physics classes. Time is put on the $x$-axis and distance $x$ is put on the $y$-axis. Somewhat confusingly, at first, as one might have gotten accustomed to using values of $x$ on the $x$-axis during the maths lessons. Of course, one learns that it is less about the names of variables and axes, rather, it is a matter of which is the independent and which is the dependent variable. The independent one, in this case, time $t$ (time flies, whether we want to or not), is then laid out over the axis called $x$ (which has not much to do with the variable named $x$), and the dependent one, a variable which happened to be named $x$, is projected onto the $y$-axis.

    In this diagram, we see four particles. The fastest, $P$, is moving away in the $x$-direction (which is up, but not necessarily up into the sky, do realise that!) covering more units of $x$ than the other three after the same time $t_1$ has passed. This is why it has a steeper slope. The slowest one is the one that is not moving at all, the stationary particle $S$. It is moving in time, which is why it exists at time $t_1$, but, spatially, it does not exist at a certain amount of units of $x$ away from the origin. In fact, it exists in exactly the same place, the origin.

    Figure 7: tx-diagram. A physicist’s diagram, where distance in space, x, is projected on the x-axis and distance in time t is projected on the y-axis. Note that the faster a particle travels, the smaller the angle of its ‘line’ through space and time with the x-axis. If the particle is stationary, it only ‘travels’ through time and the angle with the x-axis is maximised at ฯ€/2. In other words, it just goes straight up.

    Well, get yourself out of the habit: turns out that professional physicists like to flip the axes when it comes to time. In other words, they project the distance variable $x$ onto the $x$-axis, while the time variable $t$ almost invariably gets to be projected onto the $y$-axis. Yes, you heard it correctly. Time goes up in a physicist’s diagram. The esteemed professor Leonard Susskind, a theoretical physicist at Stanford University, even postulated, in part jokingly, during a lecture on the principle of least action that physicists are the only type of people who do this(beginfootnote)See, for instance, https://youtu.be/3apIZCpmdls?t=1447(endfootnote). And so, we flipped our diagram as you can see in Figure 7.

    Speaking of units, usually, time is measured in seconds and distance in metres. Usually. Though, remember when you went to visit those new friends of your parents and that one of the first things they assured their hosts is that their hometown was actually not that distant and that it was just ‘a two-hour drive’? Distance, while usually measured in kilometres between two places, is now expressed in units of time. Assuming that people legally drive โ€“ from door to door โ€“ at an average speed of $100\text{ km/h}$, the distance will be around 200 km.

    Why do people like to express distance in terms of units of time sometimes? Well, in some cases, people aren’t interested in the exact amount of kilometres, but rather tend to focus on how much of our valuable time a certain activity consumes, hence, an answer in units of time makes sense.

    Astrophysicists do another interesting distance-as-time conversion when it comes to distances between galaxies, for instance. They work with visible light reaching their telescopes, and other types of radiation. Moreover, the distances they work with are ridiculously large, especially when expressed in kilometres. So, they work with light-years, which sounds like a unit of time, but denotes a certain distance. One light-year is the distance light travels in one Julian year, which is $365.25$ days. Light travels at a speed of $c=299792458\text{ ms}^{-1}$ in the vacuum. To calculate the number of seconds in a Julian year, we multiply the number of seconds in one minute times the number of minutes in one hour times the number of hours in one day times the number of days in one Julian year:

    \begin{align}
    60\text{ s} &\times 60\text{ minutes} \times 24\text{ hours} \times 365.25\text{ days} \\
    &= 31557600\text{ s}.
    \end{align}

    Using Equation \eqref{eq:x=vt} to calculate distance $x$ light travels in one Julian year, we get

    \begin{align}
    x &= vt \text{, and because }v=c\text{, we write:} \\
    x &= ct,\label{eq:x=ct} \\
    &= 299792458\text{ ms}^{-1} \times 31557600\text{ s} \\
    &= 9460730472580800\text{ m}, \\
    &= 9460730472580.800\text{ km}.
    \end{align}

    Since light travels this ridiculously large number of kilometres, it makes perfect sense for astrophysicists to use this fact to express the distance of stars and galaxies. This way, the nearest major galaxy, Andromeda, is only about $2.5$ light-years away. This is obviously more practical than $23651826181452\text{ km}$.

    So, about describing distance in terms of units of time, we learnt that

    • in the case of a ‘normal scale’ distance such as between two towns, expressing a spatial distance in units of time makes it easier to compare with the amount of time one wishes to spend on travelling โ€“ it becomes like comparing time with time;
    • in the case of larger scale distances such as between two galaxies, expressing a spatial distance in light-units of time makes it easier to handle the impractically large numbers of the original units.

    Let us go back at our diagram in Figure 7 again. The units of both axes are not the same. The $x$-axis is the distance, which is expressed in spatial units, such as metres. The $y$-axis is the time, expressed in temporal units, such as seconds. It is hard to compare the two: the units are not the same. Also, as particle physicists are usually dealing with extremely fast particles, near the speed of light, it is impractical to be using the standard units of time. So, physicists have devised a solution to both problems. Number one: what if we expressed time in units of distance? So, that is the other way around: not distance in units of time, but time in units of distance.

    To do that, we simply use the formula as expressed in Equation \eqref{eq:x=ct}: $x=ct$. In other words, if we multiply time $t$ with the speed of light $c$, we get a distance. A little analysis of units checks out. If we multiply the units of the speed of light with the unit of time, we get a unit of distance:

    \begin{equation}
    \text{m s}^{-1} \times \text{s} = \text{m s}^{-1}\text{s} = \text{m}\frac{\text{s}}{\text{s}} = \text{m}.
    \end{equation}

    Figure 8: ct. (a) The time axis is multiplied by c, so it is easier to compare time with space, i.e. time is expressed in units of distance. (b) We set c = 1 so that the so-called world line of any particle travelling at exactly the speed of light is always at an angle of ฯ€/4 with the x-axis. Or 45ยฐ, if you are into that sort of thing.

    This does not mean we magically, qualitatively, or even hypothetically transformed the time dimension into a space dimension, even though this would be a perfect device for a cool work of science-fiction, but it does mean that we now express time in units of distance. And so, we label the $y$-axis with $ct$ as is shown in Figure 8.

    Now, to tackle the second problem, where physicists work with particles whose motions approach the speed of light at distance scales smaller than an electron in the vicinity of black holes with forces greater than you would ever encounter, it is impractical to work with the ordinary distance and time units. Furthermore, they prefer to choose the units of $ct$ and $x$ such, that a ‘line’ of a photon, e.g. light, travelling through space and time, is always depicted at an angle of $\pi/4$ or $45^\circ$ with the $x$-axis. To do so, they set the speed of light to 1. So, $c=1$. What you get is a diagram as shown in Figure 8(b). Particle $P$ is a photon, thus travelling at the speed of light. So, its ‘line’ is at the exact angle of $\pi/4$ with both the $x$- and $y$-axis, i.e. the $x$ and $ct$, respectively. All the other particles thus travel at a certain ratio of $c$, i.e. a certain ratio of 1.

    All this should tell you enough to figure out how fast a particle would be going if its ‘line’ would be drawn underneath that of $P$, i.e. at an angle smaller than $\pi/4$. And even though you should also be able to figure out if this is at all possible, we will tell you now that this is not possible.

    By the way, the term ‘line’, which we use to describe the path a particle takes through space and time in our diagrams, is called a ‘world line’ as Hermann Minkowski would have wanted us to. And the diagrams of Figures ref7 and 8 are called Minkowski diagrams. They are also called spacetime diagrams, although there is a subtle difference: Minkowski diagrams are the subset of two-dimensional diagrams within the larger set of spacetime diagrams, which contains the 3D versions, and 4D, even.

    Wick rotation revisited

    Figure 9: from it to ict. (a) The Wick rotation of the real time axis ct to the imaginary time axis ict. We projected the coordinate system of Figure 8 onto ‘the floor’ to have the ct-axis then rotated to the imaginary ict-axis by multiplication by the imaginary unit i. (b) Consequently, the world line of P gets rotated onto the imaginary plane as well.

    We are almost ready to derive the Lorentz transformations. The only thing we have to do, is Wick rotate the (real) time axis into the imaginary time axis, i.e. we rotate the $ct$-axis of Figure 8. So, we do as we did in the previous section Number sets: we multiply by the imaginary unit $i$ from the number set $\mathbb{C}$, thereby rotating the time axis of $\mathbb{R}$ into the complex plane $\mathbb{C}$ to become an imaginary axis of time.

    Figure 9 offers a geometric representation of the whole operation. We projected our original coordinate system of Figure 8 onto ‘the floor’, so to speak. We left out particles $Q$, $R$, and $S$ to keep it legible. Wick-rotating the real time axis $ct$ by multiplying by the imaginary unit $i$ then yields the imaginary time axis $ict$. Automatically, the world line of $P$ rotates along into the complex plane. Note, that the spacetime coordinates of $P$ have changed a few times in this section. They were $(x_P,t_1)$, then they became $(x_P,ct_1)$, and have ended up to become $(x_P,ict_1)$. Just the way we like it.

    Deriving the Lorentz transformations

    Invariant world line in the complex plane

    Figure 10: Wick-rotated M and E. Hermann’s M and Albert’s E frames of reference rotated at an angle ฮธ relative to each other in the complex plane about their origin. We put a little square with sides marked I and II to aid us in our trigonometric calculations.

    To end up with the Lorentz transformations as formulated in Equations \eqref{eq:Lorentz t-prime} and \eqref{eq:Lorentz x-prime} by rotating two frames of reference relative to each other in the complex plane โ€“ with an imaginary time axis โ€“ we refer to Figure 10.

    In both frames, those of Hermann and Albert, a photon $P$ travels at speed $c$. By the second postulate of Einstein’s Special Relativity[3], we know that, somehow, the value for $c$, which is chosen to be 1 in our case, is the same in both frames of reference, even though one moves relative to the other, meaning the coordinates between the frames are unequal. In Figure 2, this is represented by $v$. In Figure 10, this is represented by an angle $\theta$.

    We see that the coordinates of $P$ in $\mathcal{M}$ are $(\Delta x, ic\Delta t)$. In $\mathcal{E}$, they are $(\Delta x’, ic\Delta t’)$. They are related to each other by some proportion of angle $\theta$. Before we find that relation, we repeat our finding regarding Equation \eqref{eq:interval} in the section Invariances: the quantity $(\Delta x)^2 – (c\Delta t)^2$ is invariant. In our case, it is the interval $OP$ that is invariant, despite the fact that $P$ has different coordinates. In other words, geometrically, both $\mathcal{M}$ and $\mathcal{E}$ agree on the length of the yellow world line as you can see in Figure 10. We should proceed to show this.

    Let us first write down the expressions for the invariant yellow world line $OP$ in both frames of reference:

    \begin{align}
    \text{Hermann, standing in }\mathcal{M}\text{, says: }(OP)^2 &= (\Delta x)^2 + (ic\Delta t)^2, \\
    \text{Albert, standing in }\mathcal{E}\text{, says: }(OP)^2 &= (\Delta x’)^2 + (ic\Delta t’)^2.
    \end{align}

    And, since both expressions calculate the same invariant quantity, obviously, we can write:

    \begin{equation}
    (\Delta x)^2 + (ic\Delta t)^2 = (\Delta x’)^2 + (ic\Delta t’)^2,
    \end{equation}

    which simplifies to

    \begin{equation}
    \Delta x^2 – c^2\Delta t^2 = (\Delta x’)^2 – c^2(\Delta t’)^2.\label{eq:interval in the complex plane}
    \end{equation}

    This is the result we wanted. Whether it is with imaginary time or with real time, the quantity $\Delta x^2 – c^2\Delta t^2$ remains invariant. (Recall that $i^2=(sqrt{-1})^2=-1$.) Even in the complex plane, $\mathcal{M}$ and $\mathcal{E}$ agree on the magnitude of this quantity.

    Coordinates in terms of the other coordinates

    Let us now express the coordinates of $P$ in $\mathcal{E}$, i.e. $(\Delta x’,ic\Delta t’)$, in terms of angle $\theta$ and the coordinates of $P$ in $\mathcal{M}$, i.e. $(\Delta x,ic\Delta t)$. Firstly, we deduce an expression for $\Delta x’$ using Figure 10:

    \begin{align}
    \Delta x’ &= \Delta x\cos\theta + \text{I}, \\
    \text{I} &= ic\Delta t\sin\theta, \\
    \therefore \Delta x’ &= \Delta x\cos\theta + ic\Delta t\sin\theta.\label{eq:delta x prime}
    \end{align}

    Secondly, we deduce an expression for $ic\Delta t’$:

    \begin{align}
    ic\Delta t’ &= ic\Delta t\cos\theta – \text{II}, \\
    \text{II} &= \Delta x\sin\theta, \\
    \therefore ic\Delta t’ &= ic\Delta t\cos\theta – \Delta x\sin\theta.\label{eq:icdelta t prime}
    \end{align}

    Lastly, as we want to find the relation between $\theta$ in the complex plane and $v$ in real spacetime, we forget $P$ for a moment and now write the expression for Albert himself, sitting in $O’$ of his frame $\mathcal{E}$ in terms of the coordinates of Hermann’s frame $\mathcal{M}$. In other words, how does Hermann see Albert move? Since Albert is not moving in his own frame $\mathcal{E}$, as we said earlier, after a certain amount of time $ic\Delta t$, his $\Delta x’=0$. So, by Equation \eqref{eq:delta x prime}, we write

    \begin{equation}
    \Delta x’ = \Delta x\cos\theta + ic\Delta t\sin\theta = 0.
    \end{equation}

    Working this further, we get

    \begin{align}
    ic\Delta t\sin\theta &= -\Delta x\cos\theta, \\
    \frac{sin\theta}{\cos\theta} &= -\frac{\Delta x}{ic\Delta t}, \\
    \tan\theta &= -\frac{1}{ic}\frac{\Delta x}{\Delta t}, \\
    \tan\theta &= -\frac{1}{ic}v, \\
    \tan\theta &= -\frac{v}{ic}.
    \end{align}

    To remove the imaginary unit โ€“ being a surd โ€“ from of the denominator, we multiply the right hand side with $i/i$, yielding:

    \begin{align}
    \tan\theta &= -\frac{i}{i}\frac{v}{ic}, \\
    \tan\theta &= -i\frac{v}{-c}, \\
    therefore \tan\theta &= \frac{iv}{c}.\label{eq:tan \theta}
    \end{align}

    To recapitulate, we have now obtained Equations \eqref{eq:delta x prime}, \eqref{eq:icdelta t prime}, which express the coordinates of $P$ in $\mathcal{E}$ in terms of angle $\theta$ and the coordinates of $\mathcal{M}$. Lastly, we obtained relation \eqref{eq:tan \theta} between angle $\theta$ and speed $v$ of Albert’s frame $\mathcal{E}$ as seen by Hermann in his frame $\mathcal{M}$. So, to restate, we obtained the following transformations:

    \begin{aligned}\Delta x’ &= \Delta x\cos\theta + ic\Delta t\sin\theta,&\quad\eqref{eq:delta x prime} \\ ic\Delta t’ &= ic\Delta t\cos\theta – \Delta x\sin\theta,&\quad\eqref{eq:icdelta t prime} \\tan\theta &= \frac{iv}{c}.&\quad\eqref{eq:tan \theta}\end{aligned}

    Figure 11: Triangle tan ฮธ. The geometric representation of Equation 9: an imaginary triangle with an imaginary slope tan ฮธ, where ฮ“ is the hypotenuse. Note that sin ฮธ = (iv/c)/ฮ“ and cos ฮธ = 1/ฮ“.

    The Lorentz transformations

    Note that, algebraically, it is possible to write Equation \eqref{eq:tan \theta} as

    \begin{equation}
    \tan\theta = \frac{iv/c}{1},
    \end{equation}

    which, geometrically, looks like 11. Note that $\sin\theta=(\mathrm{iv/c})/\Gamma$ and $\cos\theta=1/\Gamma$, so all we have to do now, is figure out what $\Gamma$ is. Using, again, the Pythagorean theorem:

    \begin{align}
    \Gamma^2 &= 1^2 + \left(\frac{iv}{c}\right)^2, \\
    &= 1 + \frac{-v^2}{c^2}, \\
    \therefore \Gamma &= \sqrt{1-\frac{v^2}{c^2}}.
    \end{align}

    We can now write:

    \begin{align}
    \sin\theta &= \frac{iv/c}{\sqrt{1-v^2/c^2}}, \\
    \cos\theta &= \frac{1}{\sqrt{1-v^2/c^2}}.
    \end{align}

    This is starting to look good. Moving on to substitute $\sin\theta$ and $\cos\theta$ in Equation \eqref{eq:delta x prime}, yields:

    \begin{align}
    \Delta x’ &= \Delta x \left(\frac{1}{\sqrt{1-v^2/c^2}}\right) + ic\Delta t\left(\frac{iv/c}{\sqrt{1-v^2/c^2}}\right), \\
    &= \frac{\Delta x}{\sqrt{1-v^2/c^2}} + \frac{-v\Delta t}{\sqrt{1-v^2/c^2}}, \\
    \therefore \Delta x’ &= \frac{\Delta x-v\Delta t}{\sqrt{1-v^2/c^2}}.
    \end{align}

    Since in our configuration the differences are calculated from the origin, we can leave out the $\Delta$-sign, using just the coordinates, and so we obtain

    \begin{equation}
    x’ = \frac{x-vt}{\sqrt{1-v^2/c^2}},
    \end{equation}

    which is indeed Equation \eqref{eq:Lorentz x-prime}.

    Substituting $\sin\theta$ and $\cos\theta$ in Equation \eqref{eq:icdelta t prime}, yields:
    \begin{align}
    ic\Delta t’ &= ic\Delta t\left(\frac{1}{\sqrt{1-v^2/c^2}}\right) – \Delta x\left(\frac{iv/c}{\sqrt{1-v^2/c^2}}\right), \\
    ic\Delta t’ &= \frac{ic\Delta t}{\sqrt{1-v^2/c^2}} – \frac{iv\Delta x/c}{\sqrt{1-v^2/c^2}}, \\
    \Delta t’ &= \frac{\Delta t}{\sqrt{1-v^2/c^2}} – \frac{v\Delta x/c^2}{\sqrt{1-v^2/c^2}}, \\
    \therefore \Delta t’ &= \frac{\Delta t-v\Delta x/c^2}{\sqrt{1-v^2/c^2}}.\end{align}

    And so, leaving out the $\Delta$-sign, using just the coordinates, we obtain
    \begin{equation}
    t’ = \frac{t-vx/c^2}{\sqrt{1-v^2/c^2}},
    \end{equation}

    which is, indeed, Equation \eqref{eq:Lorentz t-prime}.

    It is important to note that, while not unusual to leave out the $\Delta$-sign, formally, it is incorrect: in Special Relativity there is no preferred (fixed) origin, hence, it is always about differences.

    Lastly, we reiterate that the term $1/\sqrt{1-v^2/c^2}$ is often written as $\gamma$ and is called the Lorentz factor. Also, in some texts, the term $v/c$ is replaced by symbol $\beta$, yielding the following equivalent expressions of the Lorentz transformations:

    \begin{align}
    ct’ &= \gamma(ct-\beta x), \\
    x’ &= \gamma(x-\beta ct), \\
    y’ &= y, \\
    z’ &= z.
    \end{align}

    Thanks to the imagination of many mathematicians and physicists before us, our ability to investigate, analyse, and calculate has become as supple and malleable as is, indeed, the fabric of the cosmos.

    Featured image: arielrobin

    [1] Hawking, S. (2001) The universe in a nutshell. New York: Bantam Books.
    [2] Walter, S. (2014) Poincarรฉ on clocks in motion. Amsterdam, Ne.
    [3] Einstein, A. (1905) โ€œZur Elektrodynamik Bewegter Kรถrper,โ€ Annalen der Physik, 322(10), pp. 891โ€“921. doi: 10.1002/andp.19053221004.
    [4] Bailey, D. H. and Borwein, J. M. (2016) Pi : the next generation : a sourcebook on the recent history of pi and its computation. Switzerland: Springer. doi: 10.1007/978-3-319-32377-0.
    [5] Cooke, R. (2005) The history of mathematics : a brief course. 2nd edn. New York, N.Y.: Wiley.
    [6] Caparrini S. (2006) On the Common Origin of Some of the Works on the Geometrical Interpretation of Complex Numbers. In: Williams K. (eds) Two Cultures. Birkhรคuser Basel, pp. 139-151.
    [7] Wick, G.C. (1954) Properties of Bethe-Salpeter Wave Functions. Physical Review, 96(4), pp. 1124-1134.

  • Real eigenvalues and eigenvectors of 3×3 matrices, example 3

    Real eigenvalues and eigenvectors of 3×3 matrices, example 3

    In these examples, the eigenvalues of matrices will turn out to be real values. In other words, the eigenvalues and eigenvectors are in $\mathbb{R}^n$.

    Download PDF

    Suppose, we have the following matrix: \begin{equation*} \mathbf{A}= \begin{pmatrix} \phantom{-}5 & 2 & 0 \\ \phantom{-}2 & 5 & 0 \\ -3 & 4 & 6 \end{pmatrix}. \end{equation*} The objective is to find the eigenvalues and the corresponding eigenvectors. 1. Characteristic equation Firstly, formulate the characteristic equation and solve it. The solutions are the eigenvalues of matrix $ \mathbf{A} $. If $ \mathbf{I} $ is the identity matrix of $ \mathbf{A} $ and $ \lambda $ is the unknown eigenvalue (represent the unknown eigenvalues), then the characteristic equation is \begin{equation*} \det(\mathbf{A}-\lambda \mathbf{I})=0. \end{equation*} Written in matrix form, we get \begin{equation} \label{eq:characteristic1} \begin{vmatrix} \phantom{-}5-\lambda & 2 & 0 \\ \phantom{-}2 & 5-\lambda & 0 \\ -3 & 4 & 6-\lambda \end{vmatrix} =0. \end{equation} Choose the row or column which are easiest to use to find the determinant of the matrix in equation \eqref{eq:characteristic1}. In other words, in this case, we will go down the last column as it contains the most zeros. This will decrease the length of the characteristic equation considerably: \begin{align*} &0\begin{vmatrix}\phantom{-}2 & 5-\lambda \\ -3 & 4 \end{vmatrix} \\ &\quad – 0\begin{vmatrix}5-\lambda & 2 \\ -3 & 4 \end{vmatrix} \\ &\qquad+ (6-\lambda)\begin{vmatrix}5-\lambda & 2 \\ 2 & 5-\lambda \end{vmatrix} = 0, \end{align*} and so \begin{equation} (6-\lambda)\begin{vmatrix}5-\lambda & 2 \\ 2 & 5-\lambda \end{vmatrix} = 0. \end{equation} Simplifying further, gives \begin{align*} (6-\lambda)[(5-\lambda)(5-\lambda)-2^2]  &= 0, \\ (6-\lambda)[25-5\lambda-5\lambda+\lambda^2-4] &= 0, \\ (6-\lambda)[\lambda^2-10\lambda+21] &= 0, \end{align*} enabling us to write a manageable form of the characteristic equation: \begin{equation} (6-\lambda)[(\lambda-3)(\lambda-7)] = 0, \end{equation} of which the solutions of $ \lambda $ are now apparent immediately: \begin{equation} \therefore \lambda = 6 \vee \lambda = 3 \vee \lambda = 7. \end{equation} Lastly, there are two ways to verify if the found values are correct. For instance, the sum of these values have to be equal to the trace of $\mathbf{A}$, which is the sum of the main diagonal of $\mathbf{A}$: \begin{equation*} \text{tr }\mathbf{A}=5+5+6=16. \end{equation*} And, indeed, the sum of the values of $\lambda$ is equal to 16 as well. We can also verify whether the product of the values of $\lambda$ is equal to $\det\mathbf{A}$: \begin{align*} \det\mathbf{A} &= 0\begin{vmatrix}\phantom{-}2&5\\-3&4\end{vmatrix} \\ &\qquad -0\begin{vmatrix}\phantom{-}5&2\\-3&4\end{vmatrix} \\ &\qquad\quad +6\begin{vmatrix}5&2\\2&5\end{vmatrix} \\ &= 6\begin{vmatrix}5&2\\2&5\end{vmatrix} \\ &= 6(5^2-2^2) \\ &= 126. \end{align*} And, indeed, the product of the values of $\lambda$ is equal to $6\cdot3\cdot7=126$. 2. Specify the eigenvalues The eigenvalues of matrix $ \mathbf{A} $ are thus $ \lambda = 6 $, $ \lambda = 3 $, and $ \lambda = 7$. 3. Eigenvector equations We rewrite the characteristic equation in matrix form to a system of three linear equations. As it is intended to find one or more eigenvectors $ \mathbf{v} $, let \begin{equation} \label{eq:v01} \mathbf{v} = \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} \end{equation} and \begin{equation} (\mathbf{A}-\lambda\mathbf{I})\mathbf{v}=\mathbf{0}. \end{equation} In which case, we can write \begin{equation} \label{eq:A-lambda I times v = 0} \begin{pmatrix} \phantom{-}5-\lambda & 2 & 0 \\ \phantom{-}2 & 5-\lambda & 0 \\ -3 & 4 & 6-\lambda \end{pmatrix} \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} = \mathbf{0}, \end{equation} which we can then write as a system of linear equations: \begin{equation*} \left\{ \begin{matrix} (5-\lambda)x_1  + 2x_2  + 0x_3 &= 0, \\ 2x_1 + (5-\lambda)x_2 + 0x_3 &= 0, \\ -3x_1 + 4x_2 + (6-\lambda)x_3 &= 0. \end{matrix} \right. \end{equation*} Simplifying this further, we have obtained the following eigenvector equations: \begin{equation} \label{eq:eigenvectorvergelijkingen01} \left\{ \begin{matrix} (5-\lambda)x_1 + 2x_2 &= 0, \\ 2x_1 + (5-\lambda)x_2 &= 0, \\ 3x_1 – 4x_2 – (6-\lambda)x_3 &= 0. \end{matrix} \right. \end{equation} 4. Substitute every obtained eigenvalue $\boldsymbol{\lambda}$ into the eigenvector equations 4.1. Eigenvalue $ \boldsymbol{\lambda = 3} $ Let’s start with eigenvalue $ \lambda = 3 $. Substituting this into the eigenvector equations \eqref{eq:eigenvectorvergelijkingen01}, we get \begin{align*} (5-3)x_1 + 2x_2 &= 0, \\ 2x_1 + (5-3)x_2 &= 0, \\ 3x_1 – 4x_2 – (6-3)x_3 &= 0. \end{align*} We can simplify this to \begin{align*} 2x_1 + 2x_2 &= 0, \\ 2x_1 + 2x_2 &= 0, \\ 3x_1 – 4x_2 – 3x_3 &= 0. \end{align*} The first equation and second equation reduce to $ x_1 = -x_2 $. Let’s substitute $ x_1 $ in the third equation. \begin{align*} 3(-x_2) – 4x_2 – 3x_3 &= 0 \\ \therefore 7x_2 &= -3x_3. \end{align*} In other words, if $ x_2 = -3 $, then $ x_3 = 7 $, and $ x_1 = 3 $. And so, we can now fill in the values of $ \mathbf{v} $ in \eqref{eq:v01}: \begin{equation} \mathbf{v} = \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} = \begin{pmatrix} \phantom{-}3 \\ -3 \\ \phantom{-}7 \end{pmatrix}. \end{equation} In other words, an eigenvector with eigenvalue $ \lambda = 3 $ is $ \begin{pmatrix}3 & -3 & 7\end{pmatrix}^T $. (Note: we deliberately write the words ‘an eigenvector’, as, for instance, the eigenvector $ \begin{pmatrix}54 & -54 & 126\end{pmatrix}^T $ is an eigenvector with this eigenvalue too. As long as $ x_1 = -x_2 $, and $ 7x_2 = -3x_3 $, in other words, as long as the ratios between $ x_1 $, $ x_2 $, and $ x_3 $ stay constant, it is an eigenvector of this eigenvalue. However, by convention we write the lowest possible integer values.) We can check the validity of the eigenvector by calculating the inner product of $\mathbf{A}$ with the eigenvector. If all went well, the outcome will be equal to the inner product of the eigenvalue with the eigenvector, in other words, $ \mathbf{Av}=\lambda \mathbf{v} $. If we write this down in matrix notation (while, for clarity, simultaneously specifying which part is which variable), indeed, we get \begin{align*} & \underbrace{\begin{pmatrix} \phantom{-}5 & 2 & 0 \\ \phantom{-}2 & 5 & 0 \\ -3 & 4 & 6 \end{pmatrix}}_{\mathbf{A}} \underbrace{\begin{pmatrix} \phantom{-}3 \\ -3 \\ \phantom{-}7 \end{pmatrix}}_{\mathbf{v}} \\ &= \begin{pmatrix} 5\cdot3 + 2\cdot-3 + 0\cdot7 \\ 2\cdot3 + 5\cdot-3 + 0\cdot7 \\ -3\cdot3 + 4\cdot-3 + 6\cdot7 \\ \end{pmatrix} \\ &= \begin{pmatrix} \phantom{-}9 \\ -9 \\ \phantom{-}21 \end{pmatrix} = \underbrace{3}_{\lambda} \underbrace{ \begin{pmatrix} \phantom{-}3 \\ -3 \\ \phantom{-}7 \end{pmatrix}}_{\mathbf{v}}. \end{align*} 4.2. Eigenvalue $ \boldsymbol{\lambda = 6} $ In the same vein, we replace $ \lambda = 6 $ in the eigenvector equations of \eqref{eq:eigenvectorvergelijkingen01}. We then write the following three linear equations: \begin{align*} (5-6)x_1 + 2x_2 &= 0, \\ 2x_1 + (5-6)x_2 &= 0, \\ 3x_1 – 4x_2 – (6-6)x_3 &= 0. \end{align*} We can simplify further: \begin{align*} x_1 &= 2x_2, \\ 2x_1 &= x_2, \\ 3x_1 &= 4x_2. \end{align*} We see that $ x_1 = x_2 = 0 $ is the only solution to this system of simultaneous equations. As $ x_3 $ always multiplies by 0 (which is why it does not appear anymore in this system), $ x_3 $ can take any value. An eigenvector to eigenvalue $ \lambda = 6 $ is therefore simply $ \begin{pmatrix}0 & 0 & 1\end{pmatrix}^T $. Checking this using the inner product of matrix $\mathbf{A}$ with $ \begin{pmatrix}0 & 0 & 1\end{pmatrix}^T $, we get, indeed, $\lambda\mathbf{v}$: \begin{align*} &\begin{pmatrix} \phantom{-}5 & 2 & 0 \\ \phantom{-}2 & 5 & 0 \\ -3 & 4 & 6 \end{pmatrix} \begin{pmatrix} 0 \\ 0 \\ 1 \end{pmatrix} \\ &= \begin{pmatrix} \phantom{-}5\cdot0 + 2\cdot0 + 0\cdot1 \\ \phantom{-}2\cdot0 + 5\cdot0 + 0\cdot1 \\ -3\cdot0 + 4\cdot0 + 6\cdot1 \\ \end{pmatrix} \\ &= \begin{pmatrix} 0 \\ 0 \\ 6 \end{pmatrix} = 6 \begin{pmatrix} 0 \\ 0 \\ 1 \end{pmatrix}. \end{align*} 4.3. Eigenvalue $ \boldsymbol{\lambda = 7} $ Substituting $\lambda=7$, yields \begin{align*} (5-7)x_1 + 2x_2 &= 0, \\ 2x_1 + (5-7)x_2 &= 0, \\ 3x_1 – 4x_2 – (6-7)x_3 &= 0. \end{align*} Working this further: \begin{align*} x_1 &= x_2, \\ x_1 &= x_2, \\ 3x_1 – 4x_2 + x_3 &= 0. \end{align*} Substituting $ x_1 = x_2 $ into the third equation, we get \begin{align*} 3x_2 – 4x_2 &= -x_3, \\ x_2 &= x_3. \end{align*} And so, we conclude \begin{equation} x_1 = x_2 = x_3. \end{equation} An eigenvector to the eigenvalue $\lambda=7$ is, thus, $\begin{pmatrix}1 & 1 & 1\end{pmatrix}^T$. Obviously, we can check this too: \begin{align*} &\begin{pmatrix} \phantom{-}5 & 2 & 0 \\ \phantom{-}2 & 5 & 0 \\ -3 & 4 & 6 \end{pmatrix} \begin{pmatrix} 1 \\ 1 \\ 1 \end{pmatrix} \\ &= \begin{pmatrix} \phantom{-}5\cdot1 + 2\cdot1 + 0\cdot1 \\ \phantom{-}2\cdot1 + 5\cdot1 + 0\cdot1 \\ -3\cdot1 + 4\cdot1 + 6\cdot1 \\ \end{pmatrix} \\ &= \begin{pmatrix} 7 \\ 7 \\ 7 \end{pmatrix} = 7 \begin{pmatrix} 1 \\ 1 \\ 1 \end{pmatrix}, \end{align*} so, it is correct.
  • Real eigenvalues and eigenvectors of 3×3 matrices, example 2

    Real eigenvalues and eigenvectors of 3×3 matrices, example 2

    In these examples, the eigenvalues of matrices will turn out to be real values. In other words, the eigenvalues and eigenvectors are in $\mathbb{R}^n$.

    Download PDF

    Suppose, we have the following matrix: \begin{equation*} \mathbf{A}= \begin{pmatrix} 8 & 0 & -5 \\ 9 & 3 & -6 \\ 10 & 0 & -7 \end{pmatrix}. \end{equation*} The objective is to find the eigenvalues and the corresponding eigenvectors. 1. Characteristic equation Firstly, formulate the characteristic equation and solve it. The solutions are the eigenvalues of matrix $ \mathbf{A} $. If $ \mathbf{I} $ is the identity matrix of $ \mathbf{A} $ and $ \lambda $ is the unknown eigenvalue (represent the unknown eigenvalues), then the characteristic equation is \begin{equation*} \det(\mathbf{A}-\lambda \mathbf{I})=0. \end{equation*} Written in matrix form, we get \begin{equation} \label{eq:characteristic1} \begin{vmatrix} 8-\lambda & 0 & -5 \\ 9 & 3-\lambda & -6 \\ 10 & 0 & -7-\lambda \end{vmatrix} =0. \end{equation} Choose the row or column which are easiest to use to find the determinant of the matrix in equation \eqref{eq:characteristic1}. In other words, in this case, start with the first element of the second column, containing $ 3-\lambda $, as the rest of the elements in this row are zero. This will decrease the length of the characteristic equation considerably: \begin{align*} &-0\begin{vmatrix}9 & -6 \\ 10 & -7-\lambda \end{vmatrix} \\ &\quad + (3-\lambda)\begin{vmatrix}8-\lambda & -5 \\ 10 & -7-\lambda \end{vmatrix} \\ &\qquad- 0\begin{vmatrix}8-\lambda & -5 \\ 9 & -6 \end{vmatrix} = 0, \end{align*} and so \begin{equation} (3-\lambda)\begin{vmatrix}8-\lambda & -5 \\ 10 & -7-\lambda \end{vmatrix} = 0. \end{equation} Simplifying further, gives \begin{align*} (3-\lambda)[(8-\lambda)(-7-\lambda)-(-5)(10)]  &= 0, \\ (3-\lambda)[-56-8\lambda+7\lambda+\lambda^2+50] &= 0, \\ (3-\lambda)[\lambda^2-\lambda-6] &= 0, \end{align*} enabling us to write a manageable form of the characteristic equation: \begin{equation} (3-\lambda)[(\lambda+2)(\lambda-3)] = 0, \end{equation} of which the solutions of $ \lambda $ are now apparent immediately: \begin{equation} \therefore \lambda = 3 \vee \lambda = -2 \vee \lambda = 3. \end{equation} Lastly, there are two ways to verify if the found values are correct. For instance, the sum of these values have to be equal to the trace of $\mathbf{A}$, which is the sum of the main diagonal of $\mathbf{A}$: \begin{equation*} \text{tr }\mathbf{A}=8+3-7=4. \end{equation*} And, indeed, the sum of the values of $\lambda$ is equal to 4 as well. We can also verify whether the product of the values of $\lambda$ is equal to $\det\mathbf{A}$: \begin{align*} \det\mathbf{A} &= -0\begin{vmatrix}9&-6\\10&-7\end{vmatrix} \\ &\qquad +3\begin{vmatrix}8&-5\\10&-7\end{vmatrix} \\ &\qquad\quad -0\begin{vmatrix}8&-5\\9&-6\end{vmatrix} \\ &= 3\begin{vmatrix}8&-5\\10&-7\end{vmatrix} \\ &= 31 \\ &= -18. \end{align*} And, indeed, the product of the values of $\lambda$ is equal to $3\cdot-2\cdot3=-18$. 2. Specify the eigenvalues The eigenvalues of matrix $ \mathbf{A} $ are thus $ \lambda = -2 $ and $ \lambda = 3$. 3. Eigenvector equations We rewrite the characteristic equation in matrix form to a system of three linear equations. As it is intended to find one or more eigenvectors $ \mathbf{v} $, let \begin{equation} \label{eq:v01} \mathbf{v} = \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} \end{equation} and \begin{equation} (\mathbf{A}-\lambda\mathbf{I})\mathbf{v}=\mathbf{0}. \end{equation} In which case, we can write \begin{equation} \label{eq:A-lambda I times v = 0} \begin{pmatrix} 8-\lambda & 0 & -5 \\ 9 & 3-\lambda & -6 \\ 10 & 0 & -7-\lambda \end{pmatrix} \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} = \mathbf{0}, \end{equation} which we can then write as a system of linear equations: \begin{equation*} \left\{ \begin{matrix}[r] (8-\lambda)x_1  + 0x_2  – 5x_3  = 0, \\ 9x_1 + (3-\lambda)x_2 – 6x_3 = 0, \\ 10x_1 + 0x_2 + (-7-\lambda)x_3 = 0. \end{matrix} \right. \end{equation*} Simplifying this further, we have obtained the following eigenvector equations: \begin{equation} \label{eq:eigenvectorvergelijkingen01} \left\{ \begin{matrix}[r] (8-\lambda)x_1 – 5x_3 &= 0, \\ 9x_1 + (3-\lambda)x_2 – 6x_3 &= 0, \\ 10x_1 + (-7-\lambda)x_3 &= 0. \end{matrix} \right. \end{equation} 4. Substitute every obtained eigenvalue $\boldsymbol{\lambda}$ into the eigenvector equations 4.1. Eigenvalue $ \boldsymbol{\lambda = -2} $ Let’s start with eigenvalue $ \lambda = -2 $. Substituting this into the eigenvector equations \eqref{eq:eigenvectorvergelijkingen01}, we get \begin{align*} (8-(-2))x_1 – 5x_3 &= 0, \\ 9x_1 + (3-(-2))x_2 – 6x_3 &= 0, \\ 10x_1 + (-7-(-2))x_3 &= 0. \end{align*} We can simplify this to \begin{align*} 10x_1 &= 5x_3, \\ 9x_1 + 5x_2 – 6x_3 &= 0, \\ 10x_1 &= 5x_3. \end{align*} The first equation and third equation reduce to $ 2x_1 = x_3 $. Let’s substitute $ x_3 $ in the second equation. \begin{align*} 9x_1 + 5x_2 -6(2x_1) &= 0 \\ \therefore 3x_1 &= 5x_2. \end{align*} In other words, if $ x_1 = 5 $, then $ x_2 = 3 $, and $ x_3 = 10 $. And so, we can now fill in the values of $ \mathbf{v} $ in \eqref{eq:v01}: \begin{equation} \mathbf{v} = \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} = \begin{pmatrix} 5 \\ 3 \\ 10 \end{pmatrix}. \end{equation} In other words, an eigenvector with eigenvalue $ \lambda = -2 $ is $ \begin{pmatrix}5 & 3 & 10\end{pmatrix}^T $. (Note: we deliberately write the words ‘an eigenvector’, as, for instance, the eigenvector $ \begin{pmatrix}30 & 18 & 60\end{pmatrix}^T $ is an eigenvector with this eigenvalue too. As long as $ 2x_1 = x_3 $, and $ 3x_1 = 5x_2 $, in other words, as long as the ratios between $ x_1 $, $ x_2 $, and $ x_3 $ stay constant, it is an eigenvector of this eigenvalue. However, by convention we write the lowest possible integer values.) We can check the validity of the eigenvector by calculating the inner product of $\mathbf{A}$ with the eigenvector. If all went well, the outcome will be equal to the inner product of the eigenvalue with the eigenvector, in other words, $ \mathbf{Av}=\lambda \mathbf{v} $. If we write this down in matrix notation (while, for clarity, simultaneously specifying which part is which variable), indeed, we get \begin{align*} & \underbrace{\begin{pmatrix} 8 & 0 & -5 \\ 9 & 3 & -6 \\ 10 & 0 & -7 \end{pmatrix}}_{\mathbf{A}} \underbrace{\begin{pmatrix} 5 \\ 3 \\ 10 \end{pmatrix}}_{\mathbf{v}} \\ &= \begin{pmatrix} 8\cdot5 + 0\cdot3 + -5\cdot10 \\ 9\cdot5 + 3\cdot3 + -6\cdot10 \\ 10\cdot5 + 0\cdot3 + -7\cdot10 \\ \end{pmatrix} \\ &= \begin{pmatrix} -10 \\ -6 \\ -20 \end{pmatrix} = \underbrace{-2}_{\lambda} \underbrace{ \begin{pmatrix} 5 \\ 3 \\ 10 \end{pmatrix}}_{\mathbf{v}}. \end{align*} 4.2. Eigenvalue $ \boldsymbol{\lambda = 3} $ In the same vein, we replace $ \lambda = 3 $ in the eigenvector equations of \eqref{eq:eigenvectorvergelijkingen01}. We then write the following three linear equations: \begin{align*} (8-3)x_1 – 5x_3 &= 0, \\ 9x_1 + (3-3)x_2 – 6x_3 &= 0, \\ 10x_1 + (-7-3)x_3 &= 0. \end{align*} We can simplify further: \begin{align*} 5x_1 &= 5x_3, \\ 9x_1 &= 6x_3, \\ 10x_1 &= 10x_3. \end{align*} We see that $ x_1=x_3=0 $ is the only solution to this system of simultaneous equations. As $ x_2 $ always multiplies by 0 (which is why it does not appear anymore in this system), $ x_2 $ can take any value. An eigenvector to eigenvalue $ \lambda = 3 $ is therefore simply $ \begin{pmatrix}0 & 1 & 0\end{pmatrix}^T $. Checking this using the inner product of matrix $\mathbf{A}$ with $ \begin{pmatrix}0 & 1 & 0\end{pmatrix}^T $, we get, indeed, $\lambda\mathbf{v}$: \begin{align*} &\begin{pmatrix} 8 & 0 & -5 \\ 9 & 3 & -6 \\ 10 & 0 & -7 \end{pmatrix} \begin{pmatrix} 0 \\ 1 \\ 0 \end{pmatrix} \\ &= \begin{pmatrix} 8\cdot0 + 0\cdot1 + -5\cdot0 \\ 9\cdot0 + 3\cdot1 + -6\cdot0 \\ 10\cdot0 + 0\cdot1 + -7\cdot0 \\ \end{pmatrix} \\ &= \begin{pmatrix} 0 \\ 3 \\ 0 \end{pmatrix} = 3 \begin{pmatrix} 0 \\ 1 \\ 0 \end{pmatrix}, \end{align*} so, that is correct.
    1. 8)(-7)-(-5)(10[]
  • Real eigenvalues and eigenvectors of 3×3 matrices, example 1

    Real eigenvalues and eigenvectors of 3×3 matrices, example 1

    In these examples, the eigenvalues of matrices will turn out to be real values. In other words, the eigenvalues and eigenvectors are in $\mathbb{R}^n$.

    Download PDF

    Suppose, we have the following matrix: \[ \mathbf{A}= \begin{pmatrix} 5 & 0 & 0 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix}. \] The objective is to find the eigenvalues and the corresponding eigenvectors. 1. Characteristic equation Firstly, formulate the characteristic equation and solve it. The solutions are the eigenvalues of matrix $ \mathbf{A} $. If $ \mathbf{I} $ is the identity matrix of $ \mathbf{A} $ and $ \lambda $ is the unknown eigenvalue (represent the unknown eigenvalues), then the characteristic equation is \[ \det(\mathbf{A}-\lambda \mathbf{I})=0. \] Written in matrix form, we get \begin{equation} \label{eq:characteristic1} \begin{vmatrix} 5-\lambda & 0 & 0 \\ 1 & 2-\lambda & 1 \\ 1 & 1 & 2-\lambda \end{vmatrix} =0. \end{equation} Choose the row or column which are easiest to use to find the determinant of the matrix in equation \eqref{eq:characteristic1}. In other words, in this case, start with the first element of the first row, $ 5-\lambda $, as the rest of the elements in this row are zero. This will decrease the length of the characteristic equation considerably: \begin{align*} &(5-\lambda)\begin{vmatrix}2-\lambda & 1 \\ 1 & 2-\lambda \end{vmatrix} \\ &\quad – 0\begin{vmatrix}1 & 1 \\ 1 & 2-\lambda \end{vmatrix} \\ &\qquad+ 0\begin{vmatrix}1 & 2-\lambda \\ 1 & 1 \end{vmatrix} = 0, \end{align*} and so \begin{equation} (5-\lambda)\begin{vmatrix}2-\lambda & 1 \\ 1 & 2-\lambda \end{vmatrix} = 0. \end{equation} Simplifying further, gives \begin{align*} (5-\lambda)[(2-\lambda)(2-\lambda)-(1)(1)]  &= 0, \\ (5-\lambda)[4-2\lambda-2\lambda+\lambda^2-1] &= 0, \\ (5-\lambda)[\lambda^2-4\lambda+3] &= 0, \end{align*} enabling us to write a manageable form of the characteristic equation: \begin{equation} (5-\lambda)[(\lambda-1)(\lambda-3)] = 0, \end{equation} of which the solutions of $ \lambda $ are now apparent immediately: \begin{equation} \therefore \lambda = 5 \vee \lambda = 1 \vee \lambda = 3. \end{equation} Lastly, there are two ways to verify if the found values are correct. For instance, the sum of these values have to be equal to the trace of $\mathbf{A}$, which is the sum of the main diagonal of $\mathbf{A}$: \begin{equation*} \text{tr }\mathbf{A}=5+2+2=9. \end{equation*} And, indeed, the sum of the values of $\lambda$ is equal to 9 as well. We can also verify whether the product of the values of $\lambda$ is equal to $\det\mathbf{A}$: \begin{align*} \det\mathbf{A} &= 5\begin{vmatrix}2&1\\1&2\end{vmatrix}-0\begin{vmatrix}1&1\\1&2\end{vmatrix}+0\begin{vmatrix}1&2\\1&1\end{vmatrix} \\ &= 5\begin{vmatrix}2&1\\1&2\end{vmatrix} \\ &= 5(2\cdot2-1\cdot1) \\ &= 15. \end{align*} And, indeed, the product of the values of $\lambda$ is equal to $5\cdot1\cdot3=15$. 2. Specify the eigenvalues The eigenvalues of matrix $ \mathbf{A} $ are thus $ \lambda = 1 $, $ \lambda = 3 $, and $ \lambda = 5 $. 3. Eigenvector equations We rewrite the characteristic equation in matrix form to a system of three linear equations. As it is intended to find one or more eigenvectors $ \mathbf{v} $, let \begin{equation} \label{eq:v01} \mathbf{v} = \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} \end{equation} and \begin{equation} (\mathbf{A}-\lambda\mathbf{I})\mathbf{v}=\mathbf{0}. \end{equation} In which case, we can write \begin{equation} \label{eq:A-lambda I times v = 0} \begin{pmatrix} 5-\lambda & 0 & 0 \\ 1 & 2-\lambda & 1 \\ 1 & 1 & 2-\lambda \end{pmatrix} \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} = \mathbf{0}, \end{equation} which we can then write as a system of linear equations: \begin{equation*} \left\{ \begin{aligned} (5-\lambda)x_1  + 0x_2  + 0x_3  &= 0, \\ x_1 + (2-\lambda)x_2 + x_3 &= 0, \\ x_1 + x_2 + (2-\lambda)x_3 &= 0. \end{aligned} \right. \end{equation*} Simplifying this further, we have obtained the following eigenvector equations: \begin{equation} \left\{ \begin{aligned} (5-\lambda)x_1 &= 0, \\ x_1 + (2-\lambda)x_2 + x_3 &= 0, \\ x_1 + x_2 + (2-\lambda)x_3 &= 0. \end{aligned} \right. \label{eq:eigenvectorvergelijkingen01} \end{equation} 4. Substitute every obtained eigenvalue $\boldsymbol{\lambda}$ into the eigenvector equations 4.1. Eigenvalue $ \boldsymbol{\lambda = 1} $ Let’s start with eigenvalue $ \lambda = 1 $. Substituting this into the eigenvector equations \eqref{eq:eigenvectorvergelijkingen01}, we get \begin{align*} (5-1)x_1 &= 0, \\ x_1 + (2-1)x_2 + x_3 &= 0, \\ x_1 + x_2 + (2-1)x_3 &= 0. \end{align*} We can simplify this to \begin{align*} 4x_1 &= 0, \\ x_1 + x_2 + x_3 &= 0, \\ x_1 + x_2 + x_3 &= 0. \end{align*} The first equation reduces to $ x_1 = 0 $  as this is obviously its only solution. The other two equations are identical. As $ x_1 = 0 $, these reduce to $ x_2 = -x_3 $, in other words, if $ x_3 = 1 $, then $ x_2 = -1 $. And so, we can now fill in the values of $ \mathbf{v} $ in \eqref{eq:v01}: \begin{equation} \mathbf{v} = \begin{pmatrix} x_1 \\ x_2 \\ x_3 \end{pmatrix} = \begin{pmatrix} \phantom{-}0 \\ -1 \\ \phantom{-}1 \end{pmatrix}. \end{equation} In other words, an eigenvector with eigenvalue $ \lambda = 1 $ is $ \begin{pmatrix}0 & -1 & 1\end{pmatrix}^T $. (Note: we deliberately write the words ‘an eigenvector’, as, for instance, the eigenvector $ \begin{pmatrix}0 & -13 & 13\end{pmatrix}^T $ is an eigenvector with this eigenvalue too. As long as $ x_2 = -x_3 $, in other words, as long as the ratio between $ x_2 $ and $ x_3 $ stays constant, it is an eigenvector of this eigenvalue. However, by convention we write the lowest possible values.) We can check the validity of the eigenvector by calculating the inner product of $\mathbf{A}$ with the eigenvector. If all went well, the outcome will be equal to the inner product of the eigenvalue with the eigenvector, in other words, $ \mathbf{Av}=\lambda \mathbf{v} $. If we write this down in matrix notation (while, for clarity, simultaneously specifying which part is which variable), indeed, we get \begin{align*} & \underbrace{\begin{pmatrix} 5 & 0 & 0 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix}}_{\mathbf{A}} \underbrace{\begin{pmatrix} \phantom{-}0 \\ -1 \\ \phantom{-}1 \end{pmatrix}}_{\mathbf{v}} \\ &= \begin{pmatrix} 5\cdot0 + 0\cdot-1 + 0\cdot1 \\ 1\cdot0 + 2\cdot-1 + 1\cdot1 \\ 1\cdot0 + 1\cdot-1 + 2\cdot1 \\ \end{pmatrix} \\ &= \begin{pmatrix} \phantom{-}0 \\ -1 \\ \phantom{-}1 \end{pmatrix} = \underbrace{1}_{\lambda} \underbrace{ \begin{pmatrix} \phantom{-}0 \\ -1 \\ \phantom{-}1 \end{pmatrix}}_{\mathbf{v}}. \end{align*} 4.2. Eigenvalue $ \boldsymbol{\lambda = 3} $ In the same vein, we replace $ \lambda = 3 $ in the eigenvector equations of \eqref{eq:eigenvectorvergelijkingen01}. We then write the following three linear equations: \begin{align*} (5-3)x_1 &= 0, \\ x_1 + (2-3)x_2 + x_3 &= 0, \\ x_1 + x_2 + (2-3)x_3 &= 0. \end{align*} We can simplify further: \begin{align*} 2x_1 &= 0, \\ x_1 – x_2 + x_3 &= 0, \\ x_1 + x_2 – x_3 &= 0. \end{align*} And here too, the first equation reduces to $ x_1 = 0 $. The second and third equation are thus both reducable to $ x_2 = x_3 $, meaning that if $x_3 = 1 $, then $ x_2 = 1 $. An eigenvector to the eigenvalue $ \lambda = 3 $, is therefore simply $ \begin{pmatrix}0 & 1 & 1\end{pmatrix}^T $. If we check this by calculating the inner product of matrix $\mathbf{A}$ with $ \begin{pmatrix}0 & 1 & 1\end{pmatrix}^T $, indeed, we obtain $\lambda\mathbf{v}$: \begin{align*} &\begin{pmatrix} 5 & 0 & 0 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix} \begin{pmatrix} 0 \\ 1 \\ 1 \end{pmatrix} \\ &= \begin{pmatrix} 5\cdot0 + 0\cdot-1 + 0\cdot1 \\ 1\cdot0 + 2\cdot-1 + 1\cdot1 \\ 1\cdot0 + 1\cdot-1 + 2\cdot1 \\ \end{pmatrix} \\ &= \begin{pmatrix} 0 \\ 3 \\ 3 \end{pmatrix} = 3 \begin{pmatrix} 0 \\ 1 \\ 1 \end{pmatrix}. \end{align*} 4.3. Eigenvalue $ \boldsymbol{\lambda = 5} $ Substituting $\lambda=5$, yields \begin{align*} (5-5)x_1 &= 0, \\ x_1 + (2-5)x_2 + x_3 &= 0, \\ x_1 + x_2 + (2-5)x_3 &= 0. \end{align*} Working this further: \begin{align*} 0x_1 &= 0, \\ x_1 – 3x_2 + x_3 &= 0, \\ x_1 + x_2 – 3x_3 &= 0. \end{align*} The first equation reduces to $0=0$, which is another way of saying that, logically, $x_1$ can take on any value as a solution to the equation. Working the other two equations some more too, we get \begin{align*} 0 &= 0, \\ x_3 &= 3x_2 – x_1, \\ x_2 &= 3x_3 – x_1. \\ \end{align*} These last two equations need some more work still. Substituting the second equation into the third, we get \begin{align*} x_2 &= 3(3x_2 – x_1)-x_1, \\ x_2 &= 9x_2 – 3x_1-x_1, \\ 8x_2 &=  4x_1, \\ x_2 &= \frac{x_1}{2}. \end{align*} Substituting this result into the second equation, gives \begin{align*} x_3 &= 3\left(\frac{x_1}{2}\right)-x_1, \\ x_3 &= \frac{3x_1}{2}-x_1, \\ x_3 &= \frac{x_1}{2}. \end{align*} And so, we conclude \begin{equation} \frac{1}{2}x_1=x_2 = x_3. \end{equation} An eigenvector to the eigenvalue $\lambda=5$ is, thus, $\begin{pmatrix}2 & 1 & 1\end{pmatrix}^T$. Obviously, we can check this too: \begin{align*} &\begin{pmatrix} 5 & 0 & 0 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix} \begin{pmatrix} 2 \\ 1 \\ 1 \end{pmatrix} \\ &= \begin{pmatrix} 5\cdot2 + 0\cdot1 + 0\cdot1 \\ 1\cdot2 + 2\cdot1 + 1\cdot1 \\ 1\cdot2 + 1\cdot1 + 2\cdot1 \\ \end{pmatrix} \\ &= \begin{pmatrix} 10 \\ 5 \\ 5 \end{pmatrix} = 5 \begin{pmatrix} 2 \\ 1 \\ 1 \end{pmatrix}, \end{align*} so, it is correct.