diff options
| author | Joshua Liu <joshua.liu@sourceobby.com> | 2026-10-04 16:47:52 -0400 |
|---|---|---|
| committer | Joshua Liu <joshua.liu@sourceobby.com> | 2026-10-04 16:47:52 -0400 |
| commit | f4213f2e2b9c2886ca25d868db7cbc12ee92bf75 (patch) | |
| tree | 48c0c7019c4a567ab30d7ad090b8e3ae080bb109 | |
| parent | 782f5ce90da374cd0198ef38bf16d88df41f51a7 (diff) | |
feat: adding information about sensors, with additional work on fault tolerant systems. Information cutoff 0923
| -rw-r--r-- | main.pdf | bin | 103006 -> 185795 bytes | |||
| -rw-r--r-- | main.tex | 95 |
2 files changed, 91 insertions, 4 deletions
| Binary files differ @@ -232,6 +232,7 @@ includefoot=true,top=19mm,nohead,footskip=12mm,bottom=6mm]{geometry} \item Energy \item Communications \end{enumerate} + \subsection{Reliability, Availability and Related Concepts} \theorybox{Definition --- Reliability}{Reliability has a specific definition for us. While we know what it MEANS to be reliable or not reliable, reliability is: \begin{center} {\bf the probability that a system will continue operation correctly during a period of time} @@ -247,17 +248,39 @@ includefoot=true,top=19mm,nohead,footskip=12mm,bottom=6mm]{geometry} A simpler example could be like scheduled maintenance. The server is down, but that is the point, maintenance is ongoing. } \theorybox{Definition --- Mean Time To Fail -- MTTF}{How long on average does it take for a system to fail?} + \subsection{Modelling Reliability} As with most things, engineers and scientists cannot resist modelling stuff with equations. Reliability is no exception. We generally say that reliability is modelled as: \begin{equation} R(t) = e^{-\lambda t} \end{equation} - where $t$ is defined as the time, and the output is the reliability of the system at $t$. So if we wanted our $.99$ reliability at $5$ hours, we would need for $\lambda$ to be somewhere around $0.002$. Ok lambda is great and all, but what does this really mean? How do we make $\lambda$ $0.002$? Can an AMD CPU be rated at $\lambda = 0.002$? So turns out, we have another, more graphical, system to help model reliability.\\ + where $t$ is defined as the time, and the output is the reliability of the system at $t$. So if we wanted our $.99$ reliability eat $5$ hours, we would need for $\lambda$ to be somewhere around $0.002$. Ok lambda is great and all, but what does this really mean? How do we make $\lambda$ $0.002$? Can an AMD CPU be rated at $\lambda = 0.002$? So turns out, we have another, more graphical, system to help model reliability. + \notebox{Why we use an exponential function to model reliability?}{We use an exponential function with a negative power because intuitively, as a system runs for a longer period of time, the less reliable it becomes.} + \theorybox{Concept --- Success Tree}{A success tree is a tree that one can build using circuit/electrical symbol diagrams that is intended to be used to model the reliability of a given system.} We can borrow from circuitry/electrical diagram design to make our model. So lets say we have a computer and a battery. We want these things to run together because they are part of the system (computer does the computer, battery provides the power, whatever). So the computer and battery are illustrated however we want. But then we connect them via an AND gate to show that they have to work together. At the outset, each component we give it some numeric value to show how reliable it is. A quality battery from Duracell might get a $0.98$, while the one you found on Temu might get a $0.20$. To model the reliability of the system with the computer and the battery, you will multiply their reliability together.\\ - This cl + This cl\\ + So that is running two things together at the same time, but what about if we can have two things, but are only used for redundancy? So one is running the other isn't. So in our previous example, we might have two batteries, one for primary power, the second as an auxiliary power source if the primary fails. For this, we use an OR gate to illustrate this. And, instead of multiplying, we take the reliability of both components, subtract it from 1 to get the probability of failure, then multiply those two together, then subtracting the product from $1$. Now our reliability goes up instead of down, genius!\\ + So then it would seem that the solution to everything is to just make everything infinitely redundant, and that way nothing ever fails. While that would be true to some extent, this is not always possible. For example if you are landing a rover on Mars, you first need to get that thing into space. Weight becomes a constraint. The heavier your lander is, the more fuel you need to put it up there. Also when you are landing your rover, you will need a bigger parachute, thus more weight. And if you are adding a redundant computer, that thing needs power too. And maybe even worse, your budget is a constraint. Maybe your job is to maximize reliability with respect to cost.\\ + \infobox{Redundancy $>$ Quality}{While having quality components is important, redundancy still trumps quality. As can be demonstrated with a success tree, having more of a slightly less reliable device compared to less of a more reliable device can yield better results.} + % \infobox{Note on modelling}{Usually people do not put a bunch of computers and batteries together. Usually they put a computer and battery together as a } + \theorybox{Concept --- Failure Tree}{A failure tree is similar to a success tree, but in the complete opposite way. But instead of giving a ``the computer has a $.005$ chance of failure'', we list what can happen, and what is the result, and how this impacts the system as a whole. So we could say ``battery is leaking'' and ``electric current comes'' would lead to ``battery catches fire'', along with failure of the computer to detect the heat will finally result in ``computer explodes''.} \subsection{How to make a system more fault tolerant} The main point of this section can be summed up as {\bf \large Redundancy}. Redundancy is the best way to achieve a more fault tolerant system. While having things such as high quality batteries, ECC memory, or well designed and implemented software, redundancy is what will let you make the most out of such equipment. - \subsubsection{Hardware} - \subsubsection{Software} + \subsubsection{Redundancy in Hardware (Components)} + As we have seen in modelling, having redundancy is key. But just how redundant should we make things? What is cost effective? + \theorybox{Concept --- Triple Modular Redundancy -- TMR}{ + TMR is the standard for anything safety critical. So ``triple'' has an obvious meaning, but what does modular mean? This means that we are referring to modules (computer $+$ battery), not components (computer, battery). + } + Although we have seen examples where modules are worse than components overall in terms of redundancy, we should remember that there are benefits to using modules.\\ + For example, imagine we need a thermostat to measure temperature in a room. (Thermostat being the module). Sure we could have a mega thermostat that never fails because it's components are all TMR, but the downside is that our reading won't be totally accurate. What if it is next to a window in the winter? Then obviously it's readings will be lower than the room's actual temperature. But if we have three thermostats, then we can take the median, or average out, the readings for higher accuracy. + \infobox{Examples of Redundancy in Hardware}{ + \begin{enumerate} + \item RAID\\ + RAID is where you have multiple drives stripped together for redundancy. But this does come with the caveat of WHAT kind of RAID you are using. RAID 0 for example, does NOT offer any redundancy benefit because of how it works. + \item Power Supplies\\ + Its common for people to have UPS' in case their PSU poops out. + \end{enumerate} + } + \subsubsection{Redundancy in Software} Software is an interesting case, since as long as the hardware doesn't give way, software will do exactly what you tell it to do, that and only that. So much thought must go into how you make the software and verify it.\\ However, when you are running software, redundancy is still important. The way that redundancy in software is achieved is by running multiple copies of the software. The way that this can achieve redundancy is that if there is some mission critical variable (such as angle of attack!), after each program has calculated its value for such variable, they can vote (especially if it is discrete), or take a mean, median, mode, whatever. If they take all the values and average them out, hopefully they will smooth out any error and come to as close of an accurate value as possible. \begin{center} @@ -265,6 +288,70 @@ includefoot=true,top=19mm,nohead,footskip=12mm,bottom=6mm]{geometry} \end{center} Given some set of requirements designed by engineers, you could have multiple teams implement the same set of requirements. By having this, your goal is to minimize the amount of faults that can occur through bugs in implementation. \theorybox{Definition --- $N$ Version Programming}{When you have $N$ teams implement the same set of requirements.} + \begin{center} + {\bf \large Good Engineering Practices} + \end{center} + TO ensure the best product, requirements for the software should be fixed. Additionally, they should be precise and detailed as to leave minimal room for architecture related bugs. + \subsubsection{Redundancy in Information} + Redundancy in information is also interesting. If we took a car for example, if the speedometer broke down, are you completely without any way of telling how fast you are going? No! + \theorybox{Idea --- You are never without information}{There is almost always some secondary source of the information you are looking for.} + If your speedometer broke down (assuming you had a good reading of the last valid speed you were going at) you can estimate based on the car's acceleration or deceleration, you could use your engine RPM, you could backwards calculate based on the distance you have travelled, or you could even use your phone (GPS, maybe the phone has it's own accelerometer to tell speed). Having redundancy in information is being able to take these secondary, correlated, sources, and correctly read their data and extract the data that you want. + \theorybox{Idea --- History Matters}{In the car example, calculating speed by making use of the last known valid speed is an example of the concept that history matters.} + \warningbox{Beware the Error!}{Although good for when the main sensor that measures the quantity that you want fails, obviously it is still best to use the main sensor when possible. Doing the calculations based on other sensors have constraints. Maybe they can only tell you part of the picture. Maybe the calculation is too slow and to get the answer in a reasonable amount of time you have to approximate. Whatever the reason is, using secondary sensors introduces noise and error into your calculations.} + % Having multiple ways to derive the quantity that you need is also beneficial in the way that \section{Sensors} + \subsection{Introduction} + This section will explain, what is a sensor, what kind of sensors are there, why sensors have noise and most importantly, how to correct for that noise. + \theorybox{Definition --- Sensor}{A sensor is some kind of device that is used to measure some physical quantity in the real world and digitize that value for a computing machine to use.} + There are many kinds of sensors, but here are some examples: + \begin{enumerate} + \item LIDAR -- Laser range finder, calculates time of flight of the laser to gage distance.\\ + \item Ultra-sound and Sonar -- same point as LIDAR but uses sound waves.\\ + \item Light Sensor -- Photodetector, or a camera: measures reflected light.\\ + \item Temperature Sensor -- Obvious\\ + \item Sound -- Microphone -- Converts vibrations created by waves of different air densities to electrical signal.\\ + \item Gyroscope -- Measures rotation\\ + \item Accelerometer -- Measures acceleration\\ + \item Chemical -- measure properties like acidity, check for specific compounds or electric noses\\ + \item Pressure Sensors -- Detect force\\ + \item Magnetic Sensors -- Detect magnetic fields such as metal detectors or compasses\\ + \item Location -- Something like GPS.\\ + \item Altitute -- Air pressure checker or barometer\\ + \item Humidity\\ + \item Electricity -- Measure current, voltage, static electricity, etc + \end{enumerate} + Although it is a sensor's job to measure such physical quantities, no sensor is completely accurate. + \theorybox{Idea --- Sensors are noisy}{This the key concept of this section. Sensors are noisy due to a variety of reasons, and it is our job to interpret the data it provides and de-noise the best we can.} + But some sensors can be less noisy than others. This is a critical component in what makes a good sensor. Other factors include: + \begin{enumerate} + \item Reliable + \item Sensitive to the quantity being measured + \item Not sensitive to other quantities + \end{enumerate} + An example of a bad sensor for example is the PacoSensor 1000+ that measures temperature with a photo camera by checking the color of a stove top. Sure you can do that but its judging temperature through color, which is bad, we should measure as directly as possible. The PacoSensor 1000+ violates the last 2 points. It is not sensitive to the quantity being measured, and it is sensitive to other quantities.\\ +A sensor should have a simple response curve, if we graph temp with respect to the response of the sensor, we want something linear or like exponential or log. A weird curve makes it harder to interpret the temperature from the response. $r(t) = f(t) = at+b$. But we generally want linear. This should at least be guaranteed in the sensor's operating range also called the dynamic range.\\ + Remember, transduction is the conversion of a physical phenomenon to an electrical signal, called $s(t)$. If $f(t)$ is physical quantity, then $s(t) = f(t) + n(t)$ with $n(t)$ being noise (conversion is noisy). For a microphone, noise can include background noise, sensor distortion, electromagnetic interference, thermal noise, etc. You can account and remove some noise, but you cannot fully eliminate it.\\ + We also have to do sampling, because $f(t)$ is continuous which computers handle that poorly. Sampling is taking a subset of a continuous curve, making discrete measurements at a uniform distribution. We then only consider those discrete points. Issues such as data loss, especially important if the continuous data changes quickly. This can be mitigated with a higher sampling rate. Also between the points we assume a straight line or constant data, which introduces distortion.\\ + Definition -- Fourier analysis is the deconstruction of a wave function into sine and cosine functions.\\ + So we can apply Fourier analysis onto $s(t)$ to break it down into sine and cosine functions with different frequencies. The point of doing this is the find the highest frequency component that matters, (which sin/cos function contributes the most to the overall wave) because sampling has to be double the highest frequency. Otherwise you lose too much data. Human sound is around 20khz, so a microphone needs to sample at 40khz. 44.1khz is the CD sampling quality. Finding the most significant frequency also helps with filtering noise.\\ + For us, we will look at Fourier analysis through the lens of discrete linear algebra. For this, ortho-normal basis are important.\\ + So how does this affect us? Suppose we get 10 samples for $s(t)$. So our result is a 1x10 vector, a point in 10-d space. We can then break down our array into a linear combination of an ortho-normal basis of 10-d space. We want a useful basis, so when choosing our basis, we take our most significant sine wave, take 10 samples in the first cycle, and use that as our first vector $\hat{f}_{s1}$, then we take $\hat{f}_{s2}$ with most significant cosine wave. This is orthogonal because of properties of the sine and cosine. After that, we do 10 samples with 2 cycles in the sine wave, so on and so forth until we have 10 vectors.\\ + Thus we can compute $s(k)$ with a linear combination of our Fourier basis. All that is left to determine is how much to scale each basis, we can compute that by doing $f_{k_i} \cdot s(k)$\\ + In real life engineers use $e^{\ell\theta}=\cos\theta + i\sin\theta$.\\ + We need to do this because the more ``square wave'' our thing looks, we need more sine or cosine waves, so we just take the most significant wave. We want to preserve our original signal, but that is impractical because of how much data we need, but we can get close enough only taking the most significant waves, this is an important concept in compression.\\ + Fourier applies to all functions. The fewer waves you use, the smoother the recreated line is.\\\\ + After we sample, we need to quantize the value into a storage value, in other words, how do we represent the data? Int, double, char?The main issue here is that we can lose information due to precision issues (caused by storage limitations and data rate or how much data you need to push, an example is how many pixels need to be set for 30 fps 4k video), thus an unavoidable introduction of distortion. This kind of noise is very annoying because this kind of noise is not random, while the techniques we will learn deal with random noise. Structured noise is in general very had to remove. If we skip quantization, we don't get this structured noise, but the issue is that the data is updated in real time by the signal, so we need a more non-volatile place to store the data.\\\\ + {\bf Noise reduction/de-noising --- Classical way}\\ + Best way to handle this is to make a reasonable assumption about the data. + \begin{enumerate} + \item Noise is IID, key being that noise is not related to each other, or the noise is not correlated (unless its structured noise). + \item Noise is zero-mean (big)\\ + This means taking the average of many successive noise values yields something close to zero. + \end{enumerate} + In practice taking these two points, for example looking at temperature. Temp usually changes very slowly. If we take a bunch of measurements really fast and take the average, we can usually recover the temperature value because noise is zero-mean.\\\\ + Another example is night photography. Take multiple shots, then average them to eliminate noise.\\ + This stops working when the signal changes too quickly. We need to sample at a rate faster than a signal is changing to use the average technique.\\ + If you cannot average, you can try to smooth out your data using Fourier transformation. But if you need to work in real time, you can define a window of time (current time minus some fixed time) and take a weighted average using Gaussian weights. But the weighted average is an imperfect noise reduction method, and it also leads to signal loss/distortion (over-smoothing)\\ + Something to note is that this only works for real values. But what about categorical sensors, where the signal output is some kind of discrete value or categorical value. Here, noise looks like the category returned keeps changing or reporting the wrong category. You can deal with this noise by taking the mode. \end{document} |
