GAO Summary: Cost Overruns, Schedule Delays, Ongoing Technical Problems With Mars Science Laboratory
“The project has additional concerns regarding the spacecraft’s software that enable its functionality once it arrives at the landing site. Project officials stated that the basic software for landing and traversing exists, but it needs to be upgraded in order to achieve full capability. The project plans to release updates and test its flight software for entry, descent, and landing (EDL) and software for surface operations during the spacecraft’s 9-month cruise phase to Mars.”
MSL Was Launched WIth Incomplete and Flawed Software
Comments are closed.

When will NASA and organisations in general learn to stop completely underestimating the amount of effort and risk from a software effort? NASA has had this problem since Apollo – I will concede that it was fair for Apollo because software was so new at the time. However we are now half a century later and we are still struggling with this problem. I have seen typical manager arguments for Flight projects, “It’s just software! How hard can it be?” Or “Oh, there is a problem with the hardware? The software can just compensate for the problem.” The appropriate answer isn’t just throwing dollars and manpower at the problem either. It’s real true Systems Engineering, Workable (Not only on paper to satisfy the process) requirements, and buy in from science/mission ops on realistic workable requirements early on. If honest workable requirements/specifications don’t come FIRST before a line of code must be written then you have just committed the software engineers to do their best to make up the specification in the spirit of what is wanted/required.
Haven’t I heard before that the operating software is the hardest part of any NEW air craft, space craft. Isn’t this a major reason to design general purpose vehicles that are used over and over with minor differences for each mission saving money and reducing risk?
Nox,
Too often it seems like software design and coding are the only technical jobs where you can get paid just for being close.
Steve
weathermen get paid for being wrong!
Excellent point boomersmom. They’re even farther off the mark than the software people. I guess either one is nice work if you can get it.
Steve
OK, I’ll take the bait. Software on a spacecraft can’t be “close” to being right. Have a reset during a burn or EDL or any other major event and you have a long day ahead of you along with a nice visit from an MIB. I don’t have direct insight into MSL, but software is one of the last pieces of the puzzle in building a spacecraft so it’s not terribly surprising that it may have a schedule hit. Additionally, it is one of the most non-deterministic elements of a spacecraft. Despite endless testing, there can still be some combination of commands or OS interaction or any number of things that aren’t always foreseen.
But since you have taken this tack, consider this an answer the previous thread about the RAD750 in which you lamented the lack of computing progress on-board. Added complexity is another reason why software can be delayed.
nkastro,
How do you make sure the software is perfect before launching? Just like you test procedures and integrated overall systems with simulations, you test software modules with emulations, but how do you test the final integrated software package? Some of it can be done with emulation, and some with simulation, but there will always be elements that you can’t test except “on the job,” in the real environment and under real conditions. This is true to some extent for almost any software development, but for mission critical software, like aircraft and spacecraft, you go through a much more extensive, controlled and documented development process, but at the end of the day there’s always a chance that some things aren’t right.
The source of my statement, really, was fact that software designers and programmers get paid by wages or salary, not by job completed, so at the end of each week/month they get paid but at the end of the scheduled job there may well be no “finished” product yet that they’ve produced. Software development is very much an iterative process. You do a design; you program it; it’s not quite right (it has bugs); so you modify the design; then you modify the programming; then you retest and oops!, another bug. You repeat the process until you can’t find any more bugs. But that doesn’t guarantee that you’ve caught all of the bugs, only the ones that your testing capabilities and design reviews can catch. As I said above, some things can only be tested on the job.
As a side note, I can give you a real world example of “only close to finished” from an aircraft program. One of the Airbus planes was undergoing initial flight tests. The flight control software is in two basic modules, LNAV and VNAV, which are Lateral (horizontal) and Vertical flight control. They test the LNAV first. In this one case, the first time they engaged the VNAV, it worked fine, but then they got a surprise – the software would not let the aircraft go below 3,000 feet once it had gone above 3,000. In other words, the plane’s software would not let the pilot land; nor would it let him disengage the VNAV after this problem occurred. The pilot had to let the plane run out of fuel and then land it dead stick to get down. (I wasn’t there, but I heard it from a reliable source). A software bug, despite extensive testing, so it wasn’t “finished”.
As for processing power, why does something like a reset happen? The press releases like to blame cosmic rays, but program overload is at least as likely to be the cause when you’re running your computer(s) at or near their upper end. Do you recall the Apollo 11 landing? They had a computer error immediately before landing, error 1201 if I recall correctly, immediately followed by error 1202. A ground controller frantically searched his documents and then instructed CapCom to ignore it and reset the computer in each case. Both errors basically meant that the computer had been asked to do more than it could handle and was now too confused to do anything. These errors had never come up in training because the same combination of tasks had never been used in training (although it was a logical set of tasks). So, my desire for more processing power is not so that we can run more and more tasks, but so that there is plenty of spare capability to prevent 1201 type errors. (Having said that, actually I do expect newer mission to involve increasingly more software, but adding more processors is an option, instead of yet newer hardened processors.) Also consider that with newer processors you get much better speed and much, much, much more memory to work with. One of the things that messes up older computers is swapping code and data between RAM (where it’s executed) and hard memory. The more RAM you have, the faster you run and the less chance you have of memory overrun type errors (the thing that made us all hate pre-XP Windows).
That’s the basics behind my comments. To be honest, it was as much tongue in cheek as anything. But notice that the item headline said “Incomplete and Flawed.” Software that doesn’t get used until late in the mission can be uploaded anytime during the flight, so it can incomplete at launch. This is not overly common (as far as I know), but it is certainly not new. And “Flawed” just means that there are one or more software bugs that they are aware of by haven’t yet fixed. Again, if it’s in software that isn’t used until late in the mission, they’ve got months to fix it and upload the updated software. There’s an element of calculated risk to this, but if the alternative was delaying the launch another two years, you can see why they went this way. I don’t think the same choice would be made on a manned flight.
Steve
Most of the recent NASA Mars probes have launched with incomplete software, and had their software completed during the flight. The hardware is a hard deadline for the launch window, whereas most of the software is only needed before landing. It seems odd that NASA does this, but it’s been done a bunch of times and seems to work.
Nox, you hit the nail on the head. Project managers go into the candy store with wide eyes then lie to themselves through the design/development cycle. Then they pray to the god of Sunk Costs so they don’t get cancelled.
This is not surprising. Software is always incomplete and flawed when it debuts, even on billion-dollar spacecraft. Can anyone give an example of complete and flawless software? I hope they can work this out. The last thing NASA needs is another embarrassing failure.
Why is this necessarily a problem? Exactly the same thing happened with Spirit and Opportunity as I recall.
This is a big whoopi-do. This is nothing new and is actually a common practice. Cassini had its software programmed during cruise. Deep Impact had its software programmed almost until the moment of impact. MERs had a big software problem once they landed and they were fully reprogrammed after their arrival.
As
understand it the EDL is significantly different from previous Mars landers
with its precision guided entry, skycrane, etc.. I expect that they will be
running tests and simulation runs through the cruise stage and will only load
the final EDL code when Curiosity arrives at Mars. I am curious to know how
many times Voyager has had its code updated.
I
think the mechanical stuff is very scary, particularly after the experience
with the bizarre sticky soil that Phoenix encountered. At least the software
can be modified. It will be very exciting to see if it all works in August.
Some of the instruments like CheMin the scientists have spent decades
perfecting.
Why is this necessarily a problem? Exactly the same thing happened with Spirit and Opportunity as I recall. Of course it’s not desirable.
Steve
Wasn’t this true for MER too?
Launching spacecraft without final landing code has been done before as has been stated. As also mentioned, hardware is the main long pole. By stretching this out, while the vehicle is in cruise, they can keep the software group busy. Once the vehicle lands and begins operation, I’m sure there will be other faults that need to be fixed.
All spacecraft that use computers have flawed software. After the
first shuttle launch was scrubbed, when the crew went back to Houston to
run simulations, they had a computer crash during a TAL abort. If
anyone ever says that software can’t be used until all flaws are found
are not being realistic. One of the MER vehicles had to stop operation for a few weeks I think when a flash I/O procedure was fixed. This is mostly hype.
The great thing about software is it’s soft. Nothing wrong with updating on the way, especially since the version they have on there now is ok for basic EDL.
So most here agree that this is not an issue. The headline is misleading because many missions, including some very successful ones that we love, are launched with incomplete/flawed software. This is just the type of thing that MSM would pick up from this site and run with, “another NASA boondoggle”, leaving out the part where it is SOP.
And they wondered why I wanted to fly a MacIntosh in space…….
Flew one successfully on three Shuttle flights….
And a few weeks in Low Earth Orbit nestled inside of a big vehicle is the exact same environment as a year above the magnetosphere on a small probe?? I am continually amazed about how little people around here know about interplanetary spaceflight. There are people who fly commercial computer in LEO (like Surrey), but it is much more challenging to so this in interplanetary space where the radiation is much higher. I know people at JPL, APL, and Ames who are researching how to do this… but it is hard work to figure out a way to deal with the radiation damage and the radiation-induced faults on non rad-hard computers. This is not something that will work just because you squint your eyes and wish really hard.
As far as software, it has been common practice to launch with minimal flight software at launch for decades… just what is needed for the interplanetary cruise. The software is updated in flight… several times for long missions like Cassini. Mars missions get EDL code after launch, Cassini got software for performing maneuvers at Saturn, other missions get what they need to operate their instruments, etc. The idea is you concentrate testing on what is important for the beginning of the mission before launch and then you use the time in transit to your destination to test the other stuff. I can imagine other approaches that might have benefits over this one, but they are currently non-standard.
Flight software is tested in many, many ways… from code reviews where there is a formal process of going through code line-by-line to running it on testbeds. The same thing is done to test ground procedures and command sequences. Even then there are occasionally glitches in flight. But in those cases the spacecraft are designed to fail into a safe configuration and await updates and fixes from the ground. In a few cases, like orbit insertion and EDL, there is no way to fail safe so extra effort must be expended to make such critical sequences bullet-proof.
This is part of a processes developed over 50 years of interplanetary spaceflight (well 50 years as of this coming August with the anniversary of Mariner 2). If anyone wants to know more about interplanetary spaceflight, some JPL training material for new ops people is online, and does a good job of walking through what needs to be done to leave Earth: http://www2.jpl.nasa.gov/ba…
The first TDRS satellite was known as a Single Event Upset Detector. Its dense RAM was continually having bits upset.