summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorJoshua Liu <joshua.liu@sourceobby.com>2026-10-04 09:55:51 -0400
committerJoshua Liu <joshua.liu@sourceobby.com>2026-10-04 09:55:51 -0400
commit782f5ce90da374cd0198ef38bf16d88df41f51a7 (patch)
treed24a27ab6f9988f7ee0869ed4c23bc3ff23588dd
parent0e5d36dff67ce6f416948bd573628b7be769bd77 (diff)
feat: more work done, cutoff 0916
-rw-r--r--main.tex10
1 files changed, 8 insertions, 2 deletions
diff --git a/main.tex b/main.tex
index 199d2fa..5d6f008 100644
--- a/main.tex
+++ b/main.tex
@@ -237,17 +237,23 @@ includefoot=true,top=19mm,nohead,footskip=12mm,bottom=6mm]{geometry}
{\bf the probability that a system will continue operation correctly during a period of time}
\end{center}
}
+ \warningbox{When we say correctly, we mean it!}{Note that this is for continuous reliability, if you say that system has $0.99$ percent reliability during a $5$ hour interval, there cannot be these consistent spikes in instability after the first hour, that means the system is not $0.99$ percent reliable during this time period.}
+
\theorybox{Definition --- Availability}{Availability is defined as:
\begin{center}
{\bf the probability that at a given moment, the system is working correctly and is available.}
\end{center}
+ Note that availability and reliability are separate concepts. A system can be reliable without being available. For example, take a hypothetical car. This car, when it detects a hazard is about to cause a crash, it will take controls away from the driver and autopilot the car to safety. During such a moment, the steering, gas and brakes did not fail, they are working exactly as intended. But they are not available to the driver!\\
+ A simpler example could be like scheduled maintenance. The server is down, but that is the point, maintenance is ongoing.
}
- \warningbox{When we say correctly, we mean it!}{Note that this is for continuous reliability, if you say that system has $0.99$ percent reliability during a $5$ hour interval, there cannot be these consistent spikes in instability after the first hour, that means the system is not $0.99$ percent reliable during this time period.}
+ \theorybox{Definition --- Mean Time To Fail -- MTTF}{How long on average does it take for a system to fail?}
As with most things, engineers and scientists cannot resist modelling stuff with equations. Reliability is no exception. We generally say that reliability is modelled as:
\begin{equation}
R(t) = e^{-\lambda t}
\end{equation}
- where $t$ is defined as the time, and the output is the reliability of the system at $t$. So if we wanted our $.99$ reliability at $5$ hours, we would need for
+ where $t$ is defined as the time, and the output is the reliability of the system at $t$. So if we wanted our $.99$ reliability at $5$ hours, we would need for $\lambda$ to be somewhere around $0.002$. Ok lambda is great and all, but what does this really mean? How do we make $\lambda$ $0.002$? Can an AMD CPU be rated at $\lambda = 0.002$? So turns out, we have another, more graphical, system to help model reliability.\\
+ We can borrow from circuitry/electrical diagram design to make our model. So lets say we have a computer and a battery. We want these things to run together because they are part of the system (computer does the computer, battery provides the power, whatever). So the computer and battery are illustrated however we want. But then we connect them via an AND gate to show that they have to work together. At the outset, each component we give it some numeric value to show how reliable it is. A quality battery from Duracell might get a $0.98$, while the one you found on Temu might get a $0.20$. To model the reliability of the system with the computer and the battery, you will multiply their reliability together.\\
+ This cl
\subsection{How to make a system more fault tolerant}
The main point of this section can be summed up as {\bf \large Redundancy}. Redundancy is the best way to achieve a more fault tolerant system. While having things such as high quality batteries, ECC memory, or well designed and implemented software, redundancy is what will let you make the most out of such equipment.
\subsubsection{Hardware}