Showing posts with label evolution. Show all posts
Showing posts with label evolution. Show all posts

Thursday, October 29, 2009

Evolution of Operation, Administration, and Maintenance (OA&M)

There are several functions performed in the PSTN under the common name OA&M. These functions include provisioning (that is, distributing all the necessary software to make systems available for delivering services), billing, maintenance, and ensuring the expected level of quality of service. The scope of the OA&M field is enormous—it deals with transmission facilities, switches, network databases, common channel signaling network elements, and so on. Because of its scope, referring to OA&M as a single task would be as much a generalization as referring to a universal computer application. As we show later in this section, the development of the PSTN OA&M has been evolutionary; as new pieces of equipment and new functions were added to the PSTN, new OA&M functions (and often new pieces of equipment) were created to deal with the administration of these new pieces of equipment and new functions. This development has been posing tremendous administrative problems for network operators. Many hope that as the PSTN and Internet ultimately converge into one network, the operations of the new network will be simpler than they are today.
Initially, all OA&M functions were performed by humans, but they have been progressively becoming automated. In the 1970s, each task associated with a piece of transmission or switching equipment was run by a task-specific application developed only for this task’s purpose. As a result, all applications were developed separately from one another. They had a simple text-based user interface—administrators used teletype terminals connected directly to the entities they administered.

In the 1980s, many tasks previously performed by humans had become fully automated. The applications were developed to emulate humans (up to the point of the programs exchanging text messages as they would appear on the screen of a teletype terminal). These applications, called operations support systems (OSSs), have been developed for the myriad OA&M functions. In most cases, a computer executing a particular OSS was connected to several systems (such as switches or pieces of transmission equipment) by RS-232 lines, and a crude ad hoc protocol was developed for the purpose of this OSS. Later, these computers served as concentrators, and they were in turn connected to the mainframes executing the OSSs. Often, introduction of a new OSS meant that more computers and more lines to connect the computers to the managed elements were needed.

You may ask why the common channel signaling network was not used for interconnection to operations systems. The answer is that this network was designed only for signaling and could not bear any additional—and unpredictable—load. As a matter of fact, a common network for interconnecting OSSs and managed elements has never been developed, although in the late 1980s and early 1990s there was a plan to develop such a network based on the OSI model. In some cases, X.25 was used; in others proprietary data networks developed by the manufacturers of computer equipment were used by telephone companies. A serious industry attempt to create a common mechanism to be used by all OA&M applications has resulted in a standard called Telecommunications Management Network (TMN) and, specifically, its part known as the Common Management Identification Protocol (CMIP), developed jointly by the International Organization for Standardization (ISO) and ITU-T.

We could not possibly even list all existing OA&M tasks. Instead we review one specific task called network traffic management (NTM). This task is important to the subject for the following three reasons. First, the very problem this task deals with is a good illustration of the vulnerability of the PSTN to events it has not been engineered to handle. (One such event—overload of PSTN circuits because of Internet traffic—has resulted in significant reengineering of the access to the Internet.) Second, the problems of switch overload and network overload are not peculiar to the PSTN—they exist (and are dealt with) today in data networks. Yet, the very characteristics of voice traffic are likely to create in the Internet and IP networks exactly the same problems once IP telephony takes off. Similar problems have similar solutions, so we expect the network traffic management applications to be useful in IP telephony. Third, IN and NTM often work on the same problems; it has been long recognized that they need to be integrated. The integration has not taken place in the PSTN yet, so it remains among the most important design tasks for the next-generation network.
NTM was developed to ensure quality of service (QoS) for PSTN voice calls. Traditionally, quality of service in the PSTN has been defined by factors like postdial delay or the fraction of calls blocked by one of the network switches. The QoS problem exists because it would be prohibitively expensive to build switches and networks that would allow us to interconnect all telephone users all the time. On the other hand, it is not necessary to do so, because not all people are using their telephones all the time. Studies have determined the proportion of users making their calls at any given time of the day and day of the week in a given time zone, and the PSTN has consequently been engineered to handle just as much traffic as needed. (Actually, the PSTN has been slightly overengineered to make up for potential fluctuations in traffic.) If a particular local switch is overloaded (that is, if all its trunks or interconnection facilities are busy), it is designed to block (that is, reject) calls.

Initially, the switches were designed to block calls only when they could not handle them independently. By the end of the 1970s, however, the understanding of a peculiar phenomenon observed in the Bell Telephone System—called the Mother’s Day phenomenon—resulted in a significant change in the way calls were blocked (as well as other aspects of the network operation).

Figure 1 demonstrates what happens with the toll network in peak circumstances. The network, engineered at the time to handle a maximum load of 1800 erlangs (an erlang is a unit measuring the load of the network: 1 erlang = 3600 calls x sec), was supposed to behave in response to ever increasing load just as depicted in the top line in the graph—to approach the maximum load and more or less stay there. In reality, however, the network experienced inexplicably decreasing performance way below the engineered level as the load increased. What was especially puzzling was that only a small portion of switches were overloaded at any time. Similar problems occurred during natural disasters—earthquakes and floods. (Fortunately, disasters have not occurred with great frequency.) Detailed studies produced an explanation: As the network attempted to build circuits to the switches that were overloaded, these circuits could not be used by other callers—even those whose calls would pass through or terminate at the underutilized switches. Thus, the root of the problem was that ineffective call attempts had been made that tied up the usable resources.

Figure 1: The Mother’s Day phenomenon.

The only solution was to block the ineffective call attempts. In order to determine such attempts, the network needed to collect in one place much information about the whole network. For this purpose, an NTM system was developed. The system polled the switches every now and then to determine their states; in addition, switches could themselves report certain extraordinary events (called alarms) asynchronously with polling. For example, every five minutes the NTM collects the values of attempts per circuit per hour (ACH) and connections per circuit per hour (CCH) from all switches in the network. If ACH is much higher than CCH, it is clear that ineffective attempts are being made. The NTM applications have been using artificial intelligence technology to develop the inference engines that would pinpoint network problems and suggest the necessary corrective actions, although they still rely on a human’s ability to infer the cause of any problem.

Overall, the problems may arise because of transmission facilities malfunction (as in cases when rats or moles chew up a fiber link—sharks have been known to do the same at the bottom of the ocean) or a breakdown of the common channel signaling system. In a physically healthy network, however, the problems are caused by use above the engineered level (for example, on holidays) or what is called focused overload, in which many calls are directed into the same geographical area. Not only natural disasters can cause overload. A PSTN service called televoting has been expected to do just that, and so is—for obvious reasons—the freephone service, such as 800 numbers in the United States. (Televoting has typically been used by TV and radio stations to gauge the number of viewers or listeners who are asked a question and invited to call either of the two given numbers free of charge. One of the numbers corresponds to a “yes” answer; the other to “no.” Fortunately, IN has built-in mechanisms for blocking such calls to prevent overload.)

Once the cause of the congestion in the network is detected, the NTM OSS deals with the problem by applying controls, that is, sending to switches and IN SCPs the commands that affect their operation. Such controls can be restrictive (for example, directionalization of trunks, making them available only in the direction leading from the congested switch; cancellation of alternative routes through congested switches; or blocking calls that are directed to congested areas) or expansive (for example, overflowing traffic to unusual routes in order to bypass congested areas). Although the idea of an expansive control appears strange at first glance, this type of control has been used systematically in the United States to fix congestion in the Northeast Corridor between Washington, D.C., and Boston, which often takes place between 9 and 11 o’clock in the morning. Since during this period most offices are still closed in California (which is three hours behind), it is not unusual for a call from Philadelphia to Boston to be routed through a toll switch in Oakland.

Overall, the applications of global network management (as opposed to specific protocols) have been at the center of attention in the PSTN industry. This trend continues today. The initial agent/manager paradigm on which both the Open Systems Interconnection (OSI) and Internet models are based has evolved into an agent-based approach, as described by Bieszad et al. (1999). In that paper, an (intelligent) agent is defined as computational entity “which acts on behalf of others, is autonomous, . . . and exhibits a certain degree of capabilities to learn, cooperate and move.” Most of the research on this subject comes in the form of application of artificial intelligence to network management problems. Agents communicate with each other using specially designed languages [such as Agent Communication Language (ACL)]; they also use specialized protocols [such as Contract-Net Protocol (CNP)]. As the result of the intensive research, two agent systems—Foundation for Intelligent Physical Agents (FIPA) and Mobile Agent System Interoperability Facilities (MASIF)—have been proposed. These specifications, however, are not applicable to the products and services described, for which reason they are not addressed here. Consider them, though, as an important reference to a technology in the making.

Tuesday, October 20, 2009

Evolution of Signaling

Now that we know what the voice circuit between the switches is, we can talk about how it is established. In the so-called plain old telephone service (POTS), establishing a call is routing, for once the call (for example, an end-to-end virtual circuit) is established, no routing decisions are to be made by the switches. There are three aspects to call establishment: First, a switch must understand the telephone number it receives in order to terminate the call on a line or route the call to the next switch in the chain; second, a switch must choose the appropriate circuit and let the next switch in the chain know what it is; third, the switches must test the circuit, monitor it, and finally release it at the end of the call. We will address the (quite important) concept of understanding the telephone number later. The other two circuit-related steps require that the switches exchange information. In the PSTN, this exchange is called signaling.
Initially, the signaling procedure was much closer to the original meaning of the word—the pieces of electric machinery involved were exchanging electrical signals. The human end user was (and still is) signaled with audio tones of different frequencies and durations.
As far as the switches are concerned, in the past, signaling was not unlike what our telephones do when we push the buttons to dial: switches exchanged audio signals using the very circuit (that is, trunk) over which the parties to the call were to speak. This type of signaling is called in-band signaling, and quite appropriately so, because it uses the voice band. There are quite a few problems with in-band signaling. Not only is it slow and quite annoying to the people who have to listen to meaningless tones, but also telephone users can produce the same tones the switches use and thereby deceive the network provider or disrupt the network.
To prevent fraud and also to improve efficiency, another form of signaling that would not use the voice band was needed. This could be achieved by using for signaling the frequencies that were out of the voice band (thus called out-of-band frequencies). Nevertheless, a channel in the telephone network is limited to the voice band, so there is no physical way to send frequencies beyond the voice band on such a channel. This limitation necessitated out-of-channel rather than out-of-band signaling. It was also obvious that much more information (concerning the characteristics of the circuits to be established, calling and called parties’ numbers, billing information, and so on) was required, and that this information could be stored and passed in the same form that was used for data processing. Hence (1) the information had to be encoded into a set of data structures and (2) these data structures had to be transformed over a separate data communications network. Thus, the concept of common channel signaling was born. Common channel signaling is signaling that is common to all voice channels but carried over none of them. Although it is clearly a misnomer, this type of signaling is often still called out-of-band signaling.
Let’s get back to the question of the switch understanding the telephone number. First of all, there are two types of numbers: those that actually correspond to the telephones that can be called and those that must be translated to the numbers of the first type. An example of the first type is a U.S. number +1-732-555-0137, which translates to a particular line in a particular central office (in New Jersey). An example of the second type is any U.S. number that starts with 1-800. The 800 prefix signals to the switch that the number by itself does not identify a particular switch or line (there is no 800 area code in the United States). Such a number designates a service (called toll-free in the United States or freephone in Europe) that is free to the caller but paid by the organization or person who receives calls.
Handling numbers of the first type is relatively straightforward—they end up in a switch’s routing table, where they are associated with the trunks or lines to be used in the act of establishing a call. The other (toll-free) numbers need translation. Naturally, a switch could translate the toll-free number, too, but such a solution would require tens of thousands of switches to be loaded with this information. The only feasible solution is to let a central database do the translation. The switch then needs to communicate with the database. [Note: The solution was figured out as early as 1979—see Faynberg et al. (1997) for the history.]
Another example where a database lookup is needed is implementation of local number portability (LNP). In the United States, the Telecommunications Act passed by the U.S. Congress in 1996 mandates the right of telephony service subscribers to keep their telephone numbers even when they change service providers. With that, subscribers can keep not only the numbers but also the features (such as call waiting) originally associated with the numbers. In the United States, the solutions are based on switches’ capabilities to query databases so as to locate the terminating switch when they encounter numbers marked as ported. (To be precise, this process requires two database dips—one to determine whether a dialed number is portable and the other to find the terminating switch.)
For both types of communications—out-of-band signaling among the switches and querying the database—the Bell Telephone System has designed a special data network called a common channel interoffice signaling (CCIS) network. When this network was introduced—in 1976—it was used only for out-of band signaling (hence interoffice). Thus the network served as a medium for communicating information about any trunk (channel) without being associated with that particular trunk. In other words, it was a medium common to all trunks, hence the term common channel. In the early 1980s, the network databases were connected to the network; thus signaling ceased to be strictly interoffice, and the I was taken away from the CCIS. Both the network and the concept became known as common channel signaling (CCS).
The architecture of the CCS network is depicted in Figure 1. The endpoints of the system are switches and network databases. The CCS routers are called signaling transfer points (STPs). Since all signaling has been outsourced to it, the CCS network must be as fast and as reliable as the network of the telephone switches. The reliability has been achieved through high redundancy: All STPs within the network are fully interconnected. Furthermore, each STP has a mated STP, with which it is connected through a high-speed link (C-link). Interconnection with other STPs is achieved through a backbone link (B-link). Finally, switches and databases are connected to STPs by A-links.
Figure 1: The common channel signaling (CCS) architecture.
Historically, there are two distinct types of protocols within common channel signaling: (1) interactions between the switches and databases that started as simple query/response messages for number translation and have evolved into service-independent protocols that support multiple services for IN technology; and (2) the protocols by means of which the switches exchange information necessary to establish, maintain, and tear down calls.
The CCS network has evolved through several releases and enhancements in the Bell System, and subsequently other telephone companies, which eventually resulted in multiple CCS networks. To ensure the interoperability of these networks as well as multivendor equipment interoperability in each of them, ITU-T has developed an international standard for common channel signaling. The latest release of this standard is called Signalling System No. 7 (SS No. 7).
Note that the official ITU-T abbreviation of this term is SS No. 7; however, the unofficial (but much easier to write and pronounce) term SS7 is used throughout the industry.We use the official term whenever we refer to the standard or its implementation in the network; we use SS7 when we refer to new classes of products (such as the SS7 gateway).

Saturday, October 17, 2009

Evolution of Switching

As noted, the first switch was a switching matrix (board) operated by a human. The 1890s saw the introduction of the first automatic step-by-step systems, which responded to rotary dial pulses from 1 to 10 (that is, digits 1 through 9 to 0). Cross-bar switches, which could set up a connection within a second, appeared in late 1930s. Step-by-step and cross-bar switches are examples of space-division switches; later, this technology evolved into that of time-division. A large step in switching development was made in the late 1960s as a consequence of the computer revolution. At that time computers were used for address translation and line selection. By 1980, stored program control as a real-time application running on a general-purpose computer coupled with a switch had become a norm.

At about the same time, a revolution in switching took place. Owing to the availability of digital transmission, it became possible to transmit voice in digital format. As the consequence, the switches went digital. For the detailed treatment of the subject, we recommend Bellamy (2000), but we are going to discuss it here because it is at the heart of the matter as far as the IP telephony is concerned. In a nutshell, the switching processes end-to-end voice in these four steps:

  1. A device scans in a round-robin fashion all active incoming trunks and samples the analog signal at a rate of 8000 times a second. The sampled signal is passed to the coder part of the coder/decoder device called a pulse-code modulation (PCM) codec, which outputs an 8-bit string encoding the value of the electric amplitude at the moment of the sample

  2. Output strings are fed into a frame whose length equals 8 times the number of active input lines. This frame is then passed to the time slot interchanger, which builds the output frame by reordering the original frame according to the connection table. For example, if input trunk number 3 is connected to output trunk number 5, then the contents of the 3rd byte of the input frame are inserted into the 5th byte of the output frame. (There is a limitation on the number of lines a time slot interchanger can support, which is determined solely by the speed at which it can perform, so the state of the art of computer architecture and microelectronics is constantly applied to building time slot interchangers. The line limitation is otherwise dealt with by cascading the devices into multistage units.)

  3. On outgoing digital trunk groups, the 8-bit slots are multiplexed into a transmission carrier according to its respective standard. (We will address transmission carriers in a moment.) Conversely, a digital switch accepts the incoming transmission frames from a transmission carrier and switches them as described in the previous step.

  4. At the destination switch, the decoder part of the codec translates the 8-bit strings coming on the input trunk back into electrical signals.

Note that we assumed that digital switches were toll offices (we called both incoming and outgoing circuits trunks). Indeed, initially only the toll switches on the top of the hierarchy went digital, but then digital telephony moved quickly down the hierarchy, and in the 1980s it migrated to the central offices and even PBXs. Furthermore, it has been moving to the local loop by means of the ISDN and digital subscriber line (DSL) technologies addressed further in this part.

The availability of digital transmission and switching has immediately resulted in much higher quality of voice, especially in cases where the parties to a call are separated by a long distance (information loss requires the presence of multiple regenerators, whose cumulative effect is significant distortion of analog signal, but the digital signals are fairly easy to restore—0s and 1s are typically represented by a continuum of analog values, so a relatively small change has no immediate, and therefore no cumulative, effect).

We conclude this section by listing the transmission carriers and formats. The T1 carrier multiplexes 24-voice channels represented by 8-bit samples into a 193-bit frame. (The extra bit is used as a framing code by alternating between 0 and 1.) With data rates of 8000 bits per second, the T1 frames are issued every 125 ms. The T1 data rate in the United States is thus 1.544 Mbps. (Incidentally, another carrier, called E1, which is used predominantly outside of the United States, carries thirty-two 8-bit samples in its frame.)

T1 carriers can be further multiplexed bit by bit into higher-order carriers, with extra bits added each time for synchronization:

  • Four T1 frames are multiplexed into a T2 frame (rate: 6.312 Mbps)

  • Six T2 frames are multiplexed into a T3 frame (rate: 44.736 Mbps)

  • Six T3 frames are multiplexed into a T4 frame (rate: 274.176 Mbps)

The ever increasing power of resulting pipes is depicted in Figure 1

Figure 1: The T-carrier multiplexing nomenclature.