{"id":1117,"date":"2006-10-31T12:00:00","date_gmt":"2006-10-31T12:00:00","guid":{"rendered":"http:\/\/orcldoug.com\/blog\/?p=1117"},"modified":"2006-10-31T12:00:00","modified_gmt":"2006-10-31T12:00:00","slug":"recovery-design-part-3-what-is-it-with-rac","status":"publish","type":"post","link":"http:\/\/orcldoug.com\/blog\/2006\/10\/31\/recovery-design-part-3-what-is-it-with-rac\/","title":{"rendered":"Recovery Design part 3 &#8211; What is it with RAC?"},"content":{"rendered":"<p>For the next few parts of this mini-series, I&#8217;m going to start looking at the initial high level design for the system and how well it meets the requirements. <\/p>\n<p>First of all I&#8217;m going to talk about the most critical database from the business perspective. I&#8217;ll call it the User Session database. It&#8217;s the most critical because, if it&#8217;s down, customers will lose their connections to what should be a 24x7x365 service. There are multiple app servers managing connections so if one of the app servers fails, the service is capable of restarting their sessions on another app server and retrieving their session information from the database. All very cool and reassuring &#8211; all that pretty redundancy. Erm, but if the database is unavailable, all of the sessions will eventually time-out, won&#8217;t they? Won&#8217;t all of the app servers prove worthless then? The database, therefore, is a Single Point Of Failure (remember that).<\/p>\n<p>The two most important bits of information I need at this stage are the proposed design (from <a href=\"http:\/\/18.133.199.212\/?p=1114\">the last blog<\/a>, although I noticed I&#8217;d used the lax &#8216;database instances&#8217; terminology so I&#8217;ve corrected it here and there)<\/p>\n<ul>\n<li>Two node RAC cluster supporting two seperate databases (<i>one of which is the User Session database<\/i>)<\/li>\n<\/ul>\n<p>and the true up-time requirements. This depends on how the application manages it&#8217;s connections and you often need to dig pretty deep with the vendor to arrive at a final answer. Our best answer at the moment is &#8216;less than 5 minutes&#8217;.<\/p>\n<p>So what&#8217;s wrong with this picture?<\/p>\n<p>Well, during the first meeting, I suggested that if maximum up-time is the focus, RAC might not be the best answer. I appreciate that might come as a shock to some people (it certainly did to the meeting), but consider this.<\/p>\n<p>If a given server has an expected uptime of 99.9%, then what&#8217;s the expected uptime of two of those servers in a two-node RAC cluster? Is it 100% (unbreakable)? 99.9%? 99.8%? Do we even know? Is it perhaps <i>more<\/i> likely to fail than a single node? Are we factoring in the possible failures of the underlying cluster filesystem or RAC or basic human error on a more complex configuration? My point is this, why oh why (rant approaching) do some people buy the message that RAC must mean 100% up-time? What will it protect you from? Node or Instance failure. How often do modern servers fail when balanced against how often something goes wrong because a) the software failed or b) a human being screwed up? <\/p>\n<p>However, I&#8217;d say that there&#8217;s an even bigger hole in this design. One that a careless elephant could drop through when out for an afternoon stroll. Remember I mentioned single points of failure? Well, we have two instances for this database on the RAC cluster, but, stone the crows &#8211; only one database! What happens if something goes wrong with that? How is RAC going to help us then?<\/p>\n<p>The answer I was given was &#8211; well nothing&#8217;s going to go wrong with that because it&#8217;s on high-end RAID-ed storage. Give me a break! Even if I forget <\/p>\n<ul>\n<li>The other week when a filesystem was mysteriously corrupted when some additional storage was attached to a server (and I trust the sysadmin &#8211; it was just some mysterious software problem)<\/li>\n<li>The site where the SAN kept falling over intermittently for a couple of weeks, to the bemusement of the SAN vendor and the anger of the customer. It was resolved in the end although it took a couple of weeks, so our customers would probably be a little disappointed!<\/li>\n<\/ul>\n<p>Even if I forget such things, what happens when some idiot (perhaps this idiot)  mis-types something and screws up the database? What do you think happens more often these days, user error or hardware (system) failure?<\/p>\n<p>Let me be crystal clear about this. There is only one database. Have you never had anything go wrong with the database itself? So what are you going to do when <i>that<\/i> happens. How are you going to patch the thing?<\/p>\n<p>To understand this stuff, you just need to have worked with OPS (Oracle Parallel Server) or RAC<br \/>\nor clusters in the past, but if you haven&#8217;t, here&#8217;s <a href=\"http:\/\/www.miracleas.dk\/WritingsFromMogens\/YouProbablyDontNeedRACUSVersion.pdf\">an excellent article<\/a> to<br \/>\nreinforce what I&#8217;m trying to say.\n<\/p>\n<p>RAC is nothing new. Clusters are nothing new. There are limits to what they will protect you from.<\/p>\n<p>Next time, I&#8217;ll address the single database problem.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>For the next few parts of this mini-series, I&#8217;m going to start looking at the initial high level design for the system and how well it meets the requirements. First of all I&#8217;m going to talk about the most critical database from the business perspective. I&#8217;ll call it the User Session database. It&#8217;s the most&hellip; <a class=\"more-link\" href=\"http:\/\/orcldoug.com\/blog\/2006\/10\/31\/recovery-design-part-3-what-is-it-with-rac\/\">Continue reading <span class=\"screen-reader-text\">Recovery Design part 3 &#8211; What is it with RAC?<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1117","post","type-post","status-publish","format-standard","hentry","category-uncategorized","entry"],"jetpack_featured_media_url":"","jetpack-related-posts":[{"id":1191,"url":"http:\/\/orcldoug.com\/blog\/2007\/01\/27\/dba-documentation-catalogue\/","url_meta":{"origin":1117,"position":0},"title":"DBA Documentation &#8211; Catalogue","date":"January 27, 2007","format":false,"excerpt":"Prompted by Linda's comment on a previous blog, I thought it might be worth writing a couple of postings on DBA documentation.The first thing you need is a Database Catalogue of some kind.Benefits1) Even an experienced DBA arriving on site won't know what servers exist, how to login to them,\u2026","rel":"","context":"With 11 comments","img":{"alt_text":"","src":"","width":0,"height":0},"classes":[]},{"id":866,"url":"http:\/\/orcldoug.com\/blog\/2006\/02\/08\/updates\/","url_meta":{"origin":1117,"position":1},"title":"Updates","date":"February 8, 2006","format":false,"excerpt":"Just a couple of small updates on some previous blogs.The database that was suffering from network latency problems between it and the app servers was moved down South last week. It all went very smoothly and that particular problem has been resolved. The 20+ hour job runs over night very\u2026","rel":"","context":"Similar post","img":{"alt_text":"","src":"","width":0,"height":0},"classes":[]},{"id":1207,"url":"http:\/\/orcldoug.com\/blog\/2007\/02\/17\/dst\/","url_meta":{"origin":1117,"position":2},"title":"DST","date":"February 17, 2007","format":false,"excerpt":"Three letters that I am positively sick of hearing. I remember reading about this first on Peter K's blog and - shame on me - it registered but wasn't top of my to do list. Well, it is now. Several other bloggers including Chris Foot have talked about the changes\u2026","rel":"","context":"With 5 comments","img":{"alt_text":"","src":"","width":0,"height":0},"classes":[]},{"id":893,"url":"http:\/\/orcldoug.com\/blog\/2005\/12\/19\/another-10046-success\/","url_meta":{"origin":1117,"position":3},"title":"Another 10046 Success","date":"December 19, 2005","format":false,"excerpt":"We're implementing a new packaged application at work. It includes a history import job that takes data in a flat-file and loads it into database tables. It performs some degree of data transformation but, in essence it inserts about 500,000 rows into one table and thousands in to a few\u2026","rel":"","context":"With 7 comments","img":{"alt_text":"","src":"","width":0,"height":0},"classes":[]},{"id":1429,"url":"http:\/\/orcldoug.com\/blog\/2008\/08\/20\/time-matters-db-time\/","url_meta":{"origin":1117,"position":4},"title":"Time Matters &#8211; DB Time","date":"August 20, 2008","format":false,"excerpt":"[In retrospect, the title of that first blog post might have suited the subject, but doesn't translate too well for subsequent related blog posts. That was a lack of planning or foresight on my part. These blog posts are tumbling out of my head in a fairly incoherent way. Maybe\u2026","rel":"","context":"With 13 comments","img":{"alt_text":"","src":"","width":0,"height":0},"classes":[]},{"id":1044,"url":"http:\/\/orcldoug.com\/blog\/2006\/08\/04\/tracing-session-activity-over-a-remote-database-link\/","url_meta":{"origin":1117,"position":5},"title":"Tracing session activity over a remote database link","date":"August 4, 2006","format":false,"excerpt":"Yesterday someone asked me how to trace a session that selects from a view in a remote database via a link. If they activated the trace on the local instance, they wouldn't see the bulk of the work which was happening on the remote instance - just a bunch of\u2026","rel":"","context":"With 2 comments","img":{"alt_text":"","src":"","width":0,"height":0},"classes":[]}],"_links":{"self":[{"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/posts\/1117","targetHints":{"allow":["GET"]}}],"collection":[{"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/comments?post=1117"}],"version-history":[{"count":0,"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/posts\/1117\/revisions"}],"wp:attachment":[{"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/media?parent=1117"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/categories?post=1117"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/orcldoug.com\/blog\/wp-json\/wp\/v2\/tags?post=1117"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}