Showing posts with label free tutorial. Show all posts
Showing posts with label free tutorial. Show all posts

Friday, March 28, 2008

Free tutors on File Systems and Network Attached Storage (NAS) LOCAL FILE SYSTEMS DATABASES AND JOURNALING Volume manager

Free tutors on File Systems and Network Attached Storage (NAS) LOCAL FILE SYSTEMS DATABASES AND JOURNALING

File systems form an intermediate layer between block-oriented hard disks and applications, with a volume manager often being used between the file system and the hard disk(Figure 4.1). Together, these manage the blocks of the disk and make these available to users and applications via the familiar directories and files. Disk subsystems provide block-oriented storage. For end users and for higher applications the handling of blocks addressed via cylinders, tracks and sectors is very cumbersome. File systems therefore represent an intermediate layer in the operating system that provides users with the familiar directories or folders and files and stores these on the block- oriented storage media so that they are hidden to the end users. This chapter introduces the basics of files systems and shows the role that they play in connection with storage networks. This chapter first of all describes the fundamental requirements that are imposed upon file systems (Section 4.1). Then network file systems, file servers and the Network Attached Storage (NAS) product category are introduced (Section 4.2). We will then show how shared disk file systems can achieve a significantly higher performance than classical network file systems (Section 4.3). The chapter concludes with a comparison with block-oriented storage networks (Fibre Channel SAN, iSCSI SAN) and Network Attached Storage (NAS)

LOCAL FILE SYSTEMS databases

File systems form an intermediate layer between block-oriented hard disks and applications, with a volume manager often being used between the file system and the hard disk (Figure 4.1). Together, these manage the blocks of the disk and make these available to users and applications via the familiar directories and files File systems and volume manager provide their services to numerous applications with various load profiles. This means that they are generic applications; their performance is not generally optimized for a specific application. Database systems such as DB2 or Oracle can get around the file system and manage

the blocks of the hard disk themselves (Figure 4.2). As a result, although the performance of the database can be increased, the management of the database is more difficult. In practice, therefore, database systems are usually configured to store their data in files that are managed by a file system. If more performance is required for a specific database, database administrators generally prefer to pay for higher performance hardware than to reconfigure the database to store its data directly upon the block-oriented hard disks.In addition to the basic services, modern file systems provide three functions – journaling, snapshots and dynamic file system expansion. Journaling is a mechanism that guarantees the consistency of the file system even after a system crash. To this end, the file system

Journaling

In addition to the basic services, modern file systems provide three functions – journaling, snapshots and dynamic file system expansion. Journaling is a mechanism that guaranteesthe consistency of the file system even after a system crash. To this end, the file system the blocks themselves

first of all writes every change to a log file that is invisible to applications and end users, before making the change in the filesystem itself. After a system crash the file system only has to run through the end of the log file in order to recreate the consistency of the file system. In file systems without journaling, typically older file systems like Microsoft's FAT32 file system or the UFS file system that is widespread in Unix systems, the consistency of the entire file system has to be checked after a system crash (file system check); in large file systems this can take several hours. In file systems without journaling it can therefore take several hours after a system crash – depending upon the size of the file system – before the data and thus the applications are back in operation.

Snapshots represent the same function as the instant copies function that is familiar from disk subsystems (cf. Section 2.7.1). Snapshots freeze the state of a file system at a given point in time. Applications and end users can access the frozen copy via a special path. As is the case for instant copies, the creation of the copy only takes a few seconds. Likewise, when creating a snapshot, care should be taken to ensure that the state of the frozen data is consistent.

compares instant copies and snapshots. An important advantage of snapshots is that they can be realized with any hardware. On the

other hand, instant copies within a disk subsystem place less load on the CPU and the buses of the server, thus leaving more system resources for the actual applications.

Volume manager

The volume manager is an intermediate layer within the operating system between the file system or database and the actual hard disks. The most important basic function of the volume manager is to aggregate several hard disks to form a large virtual hard Table 4.1 Snapshots are hardware-independent, however, they load the server's CPU Instant copy Snapshot Place of realization Disk subsystem File system Resource consumption Loads disk subsystem's controller and its buses Loads server's CPU and all buses Availability Depends upon disk subsystem (hardware-dependent) Depends upon file system (hardwarendependent) disk and make just this virtual hard disk visible to higher layers. Most volume managers provide the option of breaking this virtual disk back down into several smaller virtual hard disks and enlarging or reducing these (Figure 4.3). This virtualization within the volume manager makes it possible for system administrators to quickly react to changed storage requirements of applications such as databases and file systems. The volume manager can, depending upon its implementation, provide the same functions as a RAID controller (Section 2.4) or an intelligent disk subsystem (Section 2.7). As in snapshots, here too functions such as RAID, instant copies and remote mirroring are realized in a hardware-independent manner in the volume manager. Likewise, a RAID

controller or an intelligent disk subsystem can take the pressure off the resources of the server if the corresponding functions are moved to the storage devices. The realization of RAID in the volume manager loads not only on the server's CPU, but also on its buses

Thursday, March 27, 2008

KNOW MORE ABOUT VIRTUAL INTERFACES AND REMOTE DIRECT MEMORY ACCESS (RDMA)

KNOW MORE ABOUT VIRTUAL INTERFACES AND REMOTE DIRECT MEMORY ACCESS (RDMA)

With the Virtual Interface Architecture we move away from the communication between physical devices and finally approach the applications. VIA facilitates fast and efficient data exchange between applications that run on different computers. This requires an underlying network with a low latency and a low error rate. This means that the VIA can only be used over short distances, perhaps within a data centre or within a building. Originally, VIA was launched in 1997 by Compaq, Intel and Microsoft. Today it is a fixed component of InfiniBand. Furthermore, protocol mappings exist for Fibre Channel and Ethernet.

Today communication between applications is still relatively complicated. Incoming data is accepted by the network card, processed in the kernel of the operating system and finally delivered to the application. As part of this process, data is copied repeatedly from one buffer to the next. Furthermore, several process changes are necessary in the operating system. All in all this costs CPU power and places a load upon the system bus. As a result the communication throughput is reduced and its latency increased. The idea of VIA is to reduce this complexity by making the application and the network card exchange data directly with one another, bypassing the operating system. To this end, two applications initially set up a connection, the so-called Virtual Interface (VI): a common memory area is defined on both computers by means of which application and local network card exchange data (Figure 3.46). To send data the application fills the common memory area in the first computer with data. After the buffer has been filled with all data, the application announces by means of the send queue of the Virtual Interface and the so-called doorbell of the VI hardware that there is data to send. The VI hardware reads the data directly from the common memory area and transmits it to the VI hardware on the second computer. This does not inform the application until all data is available

in the common memory area. The operating system on the second computer is therefore bypassed, too. The Virtual Interface (VI) is the mechanism that makes it possible for the application (VI consumer) and the network card (VI NIC) to communicate directly with each other via common memory areas, bypassing the operating system. At a given point in time a Virtual Interface is connected with a maximum of one other Virtual Interface. Virtual Interfaces therefore only ever allow point-to-point communication with precisely one remote Virtual Interface. A VI provider consists of the underlying physical hardware (VI Network Interface Controller, VI NIC) and a device driver (kernel agent). Examples of VI NICs are VI-capable Fibre Channel host bus adapters, VI-capable Ethernet Network cards and InfiniBand host channel adapters. The VI NIC realizes the Virtual Interfaces and completion queues and transmits the data to other VI NICs. The kernel agent is the device driver of a VI NIC that is responsible for the management of Virtual Interfaces. Its duties include the generation and removal of Virtual Interfaces, the opening and closing of VI connections to remote Virtual Interfaces, memory management and error handling. In contrast to communication over the Virtual Interface and the completion queue, communication with the kernel agent is associated with the normal overhead such as process switching. This extra cost can ,however, be disregarded because after the Virtual Interface has been set up by the kernel agent all data is exchanged over the Virtual Interface. Applications and operating systems that communicate with each other via Virtual Interfaces are called VI consumers. Applications generally use a Virtual Interface via an intermediate layer such as sockets or MPI. Access to a Virtual Interface is provided by the user agent. The user agent first of all contacts the kernel agent in order to generate a Virtual Interface and to connect this to a Virtual Interface on a remote device (server, storage device). The actual data transfer then takes place via the Virtual Interface as described above, bypassing the operating system, meaning that the transfer is quick and

the load on the CPU is lessened. 108 I/O TECHNIQUES The Virtual Interface consists of a so-called work queue pair and the doorbells. The work queue pair consists of a send queue and a receive queue. The VI consumer can charge the VI NIC with the sending or receiving of data via the work queues, with the data itself being stored in a common memory area. The requests are processed asynchronously. VI consumer and VI NIC let each other know when new requests are pending or when the processing of requests has been concluded by means of a doorbell. A VI consumer can bundle the receipt of messages about completed requests from several Virtual Interfaces in a common completion queue. Work queues, doorbell and completion queue can be

used bypassing the operating system. Building upon Virtual Interfaces, the Virtual Interface Architecture defines two different communication models. It supports the old familiar model of sending and receiving messages, with the messages in this case being asynchronous sent and asynchronously received. An alternative communication model is the so-called Remote Direct Memory Access (RDMA). Using RDMA, distributed applications can read and write memory

areas of processes running on a different computer. As a result of VI, access to the remote memory takes place with low latency and a low CPU load.

Wednesday, March 26, 2008

Free Tutors on Interoperability of Fibre Channel SAN and IP STORAGE

Free Tutors on Interoperability of Fibre Channel SAN and IP STORAGE

Fibre Channel SANs are currently being successfully used in production environments. Nevertheless, interoperability is an issue with Fibre Channel SAN, as in all new cross manufacturer technologies. When discussing the interoperability of Fibre Channel SAN we must differentiate between the interoperability of the underlying Fibre Channel network layer, the interoperability of the Fibre Channel application protocols, such as FCP (SCSI over Fibre Channel) and the interoperability of the applications running on the Fibre Channel SAN. The interoperability of Fibre Channel SAN stands and falls by the interoperability of

FCP. FCP is the protocol mapping of the FC-4 layer, which maps the SCSI protocol on a Fibre Channel network (Section 3.3.8).The FCP is a complex piece of software that can only be implemented in the form ofa device driver. The implementation of hardware-like device drivers alone is a task that attracts errors as if by magic. The developers of FCP device drivers must therefore test extensively and thoroughly. Two general conditions make it more difficult to test the FCP device driver. The server initiates the data transfer by means of the SCSI protocol; the storage device only responds to the requests of the server. However, the idea of storage networks is to consolidate storage devices, i.e. for many servers to share a few large storage devices. Therefore, with storage networks a single storage device must be able to serve several parallel requests from different servers simultaneously. For example, it is typical for a server to be exchanging data with a storage device just when another server is scanning the Fibre Channel SAN for available storage devices. This situation requires end devices to be able to multitask. When testing multitasking systems the race conditions of the tasks to be performed come to bear: just a few milliseconds delay can lead to a completely different test result.The second difficulty encountered during testing is due to the large number of components that come together in a Fibre Channel SAN. Even when a single server is connected to a single storage device via a single switch, there are numerous possibilities that can-

not all be tested. If, for example, a Windows server is selected, there is still the choice between NT, 2000 and 2003, each with different service packs. Several manufacturers offer several different models of the Fibre Channel host bus adapter card in the server. If we take into account the various firmware versions for the Fibre Channel host bus adapter cards we find that we already have more than 50 combinations before we even select a switch. Companies want to use their storage network to connect servers and storage devices from various manufacturers, some of which are already present. The manufacturers of Fibre Channel components (servers, switches and storage devices) must therefore perform interoperability tests in order to guarantee that these components work with devices from third-party manufacturers. Right at the top of the priority list are those combinations that are required by most customers, because this is where the expected profit is the highest. The result of the interoperability test is a so-called support matrix. It specifies, for example, which storage device supports which server model with which operating system versions and Fibre Channel cards. Manufacturers of servers and storage devices often limit the Fibre Channel switches that can be used. Therefore, before building a Fibre Channel SAN you should carefully check whether the manufacturers in question state that they support the planned configuration. If the desired configuration is not listed, you can negotiate with the manufacturer regarding the payment of a surcharge to secure manufacturer support. Although non-supported configurations can work very well, if problems occur, you are left without support in critical situations. If

in any doubt you should therefore look for alternatives right at the planning stage. All this seems absolutely terrifying at first glance. However, manufacturers now support a number of different configurations. If the manufacturers' support matrices are taken into consideration, robust Fibre Channel SANs can now be operated. The operation of up-to-date operating systems such as Windows NT/2000, AIX, Solaris, HP-UX and Linux is particularly unproblematic. Fibre Channel SANs are based upon Fibre Channel networks. The incompatibility of the fabric and arbitrated loop topologies and the networking of fabrics and arbitrated loops

has already been discussed in Section 3.4.3. Within the fabric, the incompatibility of the Fibre Channel switches from different manufacturers should also be mentioned. At the end of 2003 we still recommend that when installing a Fibre Channel SAN only the switches and directors of a single manufacturer are used. Even though routing between switches and directors of different manufacturers may work as expected, and basic functions of the fabric topology such as aliasing, name server and zoning work well across different vendors in so-called 'compatibility modes'. But bear in mind that there is still only a very small installed base of mixed switch vendor configurations. A standard has been passed that addresses the interoperability of these basic functions, meaning that it is now just a matter of time before these basic functions work across every manufacturers' products. However, for new functions such as SAN security, inter-switch-link trunking or B-Ports, teething troubles with interoperability must once again be expected.In general, applications can be subdivided into higher applications that model and support the business processes and system-based applications such as file systems, databases and back-up systems. The system-based applications are of particular interest from the point of view of storage networks and storage management. The compatibility of network file systems such as NFS and CIFS is now taken for granted and hardly ever queried. As storage networks penetrate into the field of file systems, cross-manufacturer standards are becoming ever more important in this area too. A first offering is Network Data Management Protocol (NDMP, Section 7.9.4) for the back-up of NAS servers. Further down the road we expect also a customer demand for cross-vendor standards in the emerging field of storage virtualization (Chapter 5).

The subject of interoperability will preoccupy manufacturers and customers in the field of storage networks for a long time to come. Virtual Interface Architecture (VIA), Infini Band and Remote Direct Memory Access (RDMA) represent emerging new technologies that must also work in a cross-manufacturer manner. The same applies for Internet SCSI (iSCSI) and its variants like iSCSI Extensions over RDMA (iSER). iSCSI transmits the SCSI protocol via TCP/IP and, for example, Ethernet. Just like FCP, iSCSI has to serialize the SCSI protocol bit-by-bit and map it onto a complex network topology. Interoperability will therefore also play an important role in iSCSI.

IP STORAGE

Fibre Channel SANs are currently (2003) being successfully implemented in production environments. Nevertheless, the industry is at pains to establish storage networks based upon IP (IP storage) and Ethernet as an alternative to Fibre Channel. This section first introduces various protocols for the transmission of storage data traffic via TCP/IP (Section 3.5.1). Then we explain to what extent TCP/IP and Ethernet are suitable transmission techniques for storage networks at all (Section 3.5.2). Finally, we discuss a migration path from SCSI and Fibre Channel to IP storage (Section 3.5.3).

 

 

Tuesday, March 25, 2008

FREE TUTORS ON THE FIBRE CHANNEL PROTOCOL STACK COMMON SERVICES Link services: login and addressing and (Fabric services: name server and co)

THE FIBRE CHANNEL PROTOCOL STACK COMMON SERVICES (Fabric services: name server and co)

FC-3 has been in its conceptual phase since 1988; in currently available products FC-3 is empty. The following functions are being discussed for FC-3:

• Striping manages several paths between multiport end devices. Striping could distribute the frames of an exchange over several ports and thus increase the throughput between the two devices.

• Multi patching combines several paths between two multiport end devices to form a logical path group. Failure or overloading of a path can be hidden from the higher protocol layers.

• Compressing the data to be transmitted, preferably realized in the hardware on the host bus adapter.

• Encryption of the data to be transmitted, preferably realized in the hardware on the host bus adapter.

• Finally, mirroring and other RAID levels are the last example that are mentioned in the Fibre Channel standard as possible functions of FC-3. However, the fact that these functions are not realized within the Fibre Channel protocol does not mean that they are not available at all. For example, multipathing functions are currently provided both by suitable additional software in the operating system (Section 6.3.1) and also by some more modern Fibre Channel switches

(ISL Trucking).3.3.6 Link services: login and addressing Link services and the fabric services discussed in the next section stand next to the Fibre

Channel protocol stack. They are required to operate data traffic over a Fibre Channel network. Activities of these services do not result from the data traffic of the application protocols. Instead, these services are required to manage the infrastructure of a Fibre Channel network and thus the data traffic on the level of the application protocols. For example, at any given time the switches of a fabric know the topology of the whole network. Login Two ports have to get to know each other before application processes can exchange data over them. To this end the Fibre Channel standard provides a three-stage login mechanism (Figure 3.20): 1. Fabric login (FLOGI)The fabric login establishes a session between an N-Port and a corresponding F-Port. The fabric login takes place after the initialization of the link and is an absolute prerequisite for the exchange of further frames. The F-Port assigns the N-Port a dynamic

address. In addition, service parameters such as the buffer-to-buffer credit are negotiated. The fabric login is crucial for the point-to-point topology and for the fabric topology. An N-Port can tell from the response of the corresponding port whether it is a fabric topology or a point-to-point topology. In arbitrated loop topology the fabric login is optional. 2. N-Port login (PLOGI) N-Port login establishes a session between two N-ports. The N-Port login takes place after the fabric login and is a compulsory prerequisite for the data exchange at FC-4 level. N-Port login negotiates service parameters such as end-to-end credit. N-Port login

is optional for Class 3 communication and compulsory for all other service classes.3. Process login (PRLI) Process login establishes a session between two FC-4 processes that are based upon two different N-Ports. These could be system processes in Unix systems and system partitions in mainframes. Process login takes place after the N-Port login. Process login is optional from the point of view of FC-2. However, some FC-4 protocol mappings call for a process login for the exchange of FC-4-specific service parameters.

Addressing

Fibre Channel differentiates between addresses and names. Fibre Channel devices (servers, switches, ports) are differentiated by a 64-bit identifier. The Fibre Channel standard defines different name formats for this. Some name formats guarantee that such a 64-bit identifier will only be issued once world-wide. Such identifiers are thus also known as World Wide Names (WWPN). On the other hand, 64-bit identifiers that can be issued several times inseparate networks are simply called Fibre Channel Names (FCN).In practice this fine distinction between WWN and FCN is hardly ever noticed, with all 64-bit identifiers being called WWNs. In the following we comply with the general usage and use only the term WWN. World Wide Names are differentiated into World Wide Port Names (WWPNs) andWorld Wide Node Names (WWNNs). As the name suggests, every port is assigned its own World Wide Name in the form of a World Wide Port Name and in addition the entire device is assigned its own World Wide Name in the form of a World Wide Node Name. The differentiation between World Wide Node Name and World Wide Port Name allows us to determine which ports belong to a common multiport device in the FibreChannel network. Examples of multiport devices are intelligent disk subsystems with several Fibre Channel ports or servers with several Fibre Channel host bus adapter cards.WWNNs could also be used to realize services such as striping over several redundant physical paths within the Fibre Channel protocol. As discussed above (Section 3.3.5,

'FC-3: common services'), the Fibre Channel standard unfortunately does not support these options, so that such functions are implemented in the operating system or by manufacturer-specific expansions of the Fibre Channel standard. In the fabric, each 64-bit World Wide Port Name is automatically assigned a 24-bit port address (N-Port identifier, N-Port ID) during fabric login. The 24-bit port addresses are used within a Fibre Channel frame for the identification of transmitter and receiver of the frame. The port address of the transmitter is called the Source Identifier (S ID) and that of the receiver the Destination Identifier (D ID). The 24-bit addresses are hierarchically structured and mirror the topology of the Fibre Channel network. As a result, it is a simple matter for a Fibre Channel switch to recognize which port it must send an incoming frame to from the destination ID (Figure 3.21). Some of the 24-bit addresses are reserved

for special purposes, so that 'only' 15.5 million addresses remain for the addressing of devices. In the arbitrated loop every 64-bit World Wide Port Name is even assigned only an eight-bit address, the so-called Arbitrated Loop Physical Address (AL PA). Of the 256possible eight-bit addresses, only those for which the 8b/10b encoded transmission word contains an equal number of zeros and ones may be used. Some ordered sets for the configuration of the arbitrated loop are parametrized using AL PAs. Only by limiting thevalues for AL PAs is it possible to guarantee a uniform distribution of zeros and onesin the whole data stream. After the deduction of a few of these values for the control ..., Fibre Channel differentiates end devices using World Wide Node Names (WWPN). Each connection port is assigned its own World Wide Port Name (WWPN). For addressing in the fabric WWNNs or WWPNs are converted into shorter Port IDs that reflect the network topology of the arbitrated loop, 127 addresses of the 256 possible addresses remain. One of these addresses is reserved for a Fibre Channel switch so only 126 servers or storage devices can be connected in the arbitrated loop.

Fabric services: name server and co

In a fabric topology the switches manage a range of information that is required for the operation of the fabric. This information is managed by the so-called fabric services. All services have in common that they are addressed via FC-2 frames and can be reached by defined addresses (Table 3.3). In the following we introduce the fabric login server, the fabric controller and the name server. The fabric login server processes incoming fabric login requests under the address

'0×FF FF FE'. All switches must support the fabric login under this address. The fabric controller manages changes to the fabric under the address '0×FF FF FD'.

N-Ports can register for state changes in the fabric controller (State Change Registration, SCR). The fabric controller then informs registered N-Ports of changes to the fabric (Registered State Change Notification, RSCN). Servers can use this service to monitor their storage devices. The name server (Simple Name Server to be precise) administers a database on N-Ports under the address '0×FF FF FC'. It stores information such as port WWN, node WWN, port address, supported service classes, supported FC-4 protocols, etc. N-Ports can register



their own properties with the name server and request information on other N-Ports. Like all services, the name server appears as an N-Port to the other ports. N-Ports must log on with the name server by means of port login before they can use its services.

Monday, March 24, 2008

FREE TUTORS ON THE FIBRE CHANNEL PROTOCOL STACK FLOW CONTROL (Service classes)

Free Tutors on THE FIBRE CHANNEL PROTOCOL STACK

FC-1: 8b/10b encoding, ordered sets and link control protocolFC-1 defines how data is encoded before it is transmitted via a Fibre Channel cable(8b/10b encoding). FC-1 also describes certain transmission words (ordered sets) that are required for the administration of a Fibre Channel connection (link control protocol). 8b/10b encoding In all digital transmission techniques, transmitter and receiver must synchronize their clock-pulse rates. In parallel buses the bus rate is transmitted via an additional data line. By contrast, in the serial transmission used in Fibre Channel only one data line is available through which the data is transmitted. This means that the receiver must regenerate the transmission rate from the data stream.

The receiver can only synchronize the rate at the points where there is a signal change in the medium. In simple binary encoding (Figure 3.11) this is only the case if the signal changes from '0' to '1' or from '1' to '0'. In Manchester encoding there is a signal change for every bit transmitted. Manchester encoding therefore creates two physical signals for each bit transmitted. It therefore requires a transfer rate that is twice as high as that for binary encoding. Therefore, Fibre Channel – like many other transmission techniques – uses binary encoding, because at a given rate of signal changes more bits can be transmitted than is the case for Manchester encoding. The problem with this approach is that the signal steps that arrive at the receiver are not always the same length (jitter). This means that the signal at the receiver is sometimes a little longer and sometimes a little shorter (Figure 3.12). In the escalator analogy this means that the escalator bucks. Jitter can lead to the receiver losing synchronization with the received signal. If, for example, the transmitter sends a sequence of ten zeros, the receiver cannot decide whether it is a sequence of nine, ten or eleven zeros. If we nevertheless wish to use binary encoding, then we have to ensure that the data stream generates a signal change frequently enough that jitter cannot strike. The so-called 8b/10b encoding represents a good compromise. 8b/10b encoding converts an eight-bit

byte to be transmitted into a ten-bit character, which is sent via the medium instead of the eight-bit byte. For Fibre Channel this means, for example, that a useful transfer rate of 100 MByte/s requires a raw transmission rate of 1 Gbit/s instead of 800 Mbit/s. Incidentally, 8b/10b encoding is also used for the Enterprise System Connection Architecture (ESCON), Serial Storage Architecture (SSA), Gigabit Ethernet and InfiniBand. Finally, it should be noted that 1 Gigabyte Fibre Channel uses the 64b/66b encoding variant for a certain cable type (single lane with serial transmission). Expanding the eight-bit data bytes to ten-bit transmission character gives rise to the following advantages: • In 8b/10b encoding, of all available ten-bit characters, only those that generate a bit

sequence that contains a maximum of five zeros one after the other or five ones one after the other for any desired combination of the ten-bit character are selected. There- fore, a signal change takes place at the latest after five signal steps, so that the clock synchronization of the receiver is guaranteed.

• A bit sequence generated using 8b/10b encoding has a uniform distribution of zeros and ones. This has the advantage that only small direct currents flow in the hardware that processes the 8b/10b encoded bit sequence. This makes the realization of Fibre Channel hardware components simpler and cheaper.

• Further ten-bit characters are available that do not represent eight-bit data bytes. These additional characters can be used for the administration of a Fibre Channel link.

Ordered sets

Fibre Channel aggregates four ten-bit transmission characters to form a 40-bit transmission word. The Fibre Channel standard differentiates between two types of transmission word: data words and ordered sets. Data words represent a sequence of four eight-bit data bytes. Data words may only stand between a Start-of-Frame delimiter (SOF delimiter) and an End-of-Frame delimiter (EOF delimiter).Ordered sets may only stand between an EOF delimiter and a SOF delimiter, with SOF sand EOFs themselves being ordered sets. All ordered sets have in common that they begin with a certain transmission character, the so-called K28.5 character. The K28.5 character includes a special bit sequence that does not occur elsewhere in the data stream. The input channel of a Fibre Channel port can therefore use the K28.5 character to divide the continuous incoming bit stream into 40 bit transmission words when initializing a Fibre Channel link or after the loss of synchronization on a link. Link control protocol With the aid of ordered sets, FC-1 defines various link level protocols for the initialization and Administration of a link. The initialization of a link is the prerequisite for data exchange by means of frames. Examples of link level protocols are the initialization and arbitration of an arbitrated loop.

3.3.4 FC-2: data transfer

FC-2 is the most comprehensive layer in the Fibre Channel protocol stack. It determines how larger data units (for example, a file) are transmitted via the Fibre Channel network. It regulates the flow control that ensures that the transmitter only sends the data at a speed that the receiver can process it. And it defines various service classes that are tailored to the requirements of various applications. Exchange, sequence and frame FC-2 introduces a three-layer hierarchy for the transmission of data (Figure 3.13). At the top layer a so-called exchange defines a logical communication connection between two end devices. For example, each process that reads and writes data could be assigned its own exchange. End devices (servers and storage devices) can simultaneously maintain several exchange relationships, even between the same ports. Different exchanges help the FC-2 layer to deliver the incoming data quickly and efficiently to the orrect receiver in the higher protocol layer (FC-3). A sequence is a larger data unit that is transferred from a transmitter to a receiver. Only one sequence can be transferred after another within an exchange. FC-2 guarantees that sequences are delivered to the receiver in the same order they were sent from the transmitter; hence the name 'sequence'. Furthermore, sequences are only delivered to the next protocol layer up when all frames of the sequence have arrived at the receiver (Figure 3.13). A sequence could represent the writing of a file or an individual database transaction. A Fibre Channel network transmits control frames and data frames. Control frames contain no useful data, they signal events such as the successful delivery of a data frame. Data frames transmit up to 2112 bytes of useful data. Larger sequences therefore have to be broken down into several frames. Although it is theoretically possible to agree upon

different maximum frame sizes, this is hardly ever done in practice. A Fibre Channel frame consists of a header, useful data (payload) and a CRC checksum

(Figure 3.14). In addition, the frame is bracketed by a Start-of-Frame delimiter (SOF) and an End-of-Frame delimiter (EOF). Finally, six filling words must be transmitted by means of a link between two frames. In contrast to Ethernet and TCP/IP, Fibre Channel is an integrated whole: the layers of the Fibre Channel protocol stack are so well harmonizedwith one another that the ratio of payload to protocol overhead is very efficient at up to 98%. The CRC checking procedure is designed to recognize all transmission errors if the underlying medium does not exceed the specified error rate of 10−12 Error correction takes place at sequence level: if a frame of a sequence is wrongly transmitted, the entire sequence is retransmitted. At gigabit speed it is more efficient to resend a complete sequence than to extend the Fibre Channel hardware so that individual lost frames can be resent and inserted in the correct position. The underlying protocol

layer must maintain the specified maximum error rate of 10−12so that this procedures efficient.

Flow control

Flow control ensures that the transmitter only sends data at a speed that the receiver can receive it. Fibre Channel uses the so-called credit model for this. Each credit represents the capacity of the receiver to receive a Fibre Channel frame. If the receiver awards the transmitter a credit of '4', the transmitter may only send the receiver four frames. The transmitter may not send further frames until the receiver has acknowledged the receipt of at least some of the transmitted frames. FC-2 defines two different mechanisms for flow control: end-to-end flow control and link flow control (Figure 3.15). In end-to-end flow control two end devices negotiate the end-to-end credit before the data exchange. The end-to-end flow control is realized on the host bus adapter cards of the end devices. By contrast, link flow control takes place at each physical connection. This is achieved by two communicating ports negotiating the buffer-to-buffer credit. This means that the link flow control also takes place at the Fibre Channel switches.

Service classes

The Fibre Channel standard defines six different service classes for data exchange between end devices. Three of these defined classes (Class 1, Class 2 and Class 3) are realized in products available on the market, with hardly any products providing the connection- oriented Class 1. Almost all new Fibre Channel products (host bus adapters, switches, storage devices) support the service classes Class 2 and Class 3, which realize a packet- oriented service (datagram service). In addition, Class F serves for the data exchange between the switches within a fabric. Class 1 defines a connection-oriented communication connection between two node ports: a Class 1 connection is opened before the transmission of frames. This specifies a route through the Fibre Channel network. Thereafter, all frames take the same route through the Fibre Channel network so that frames are delivered in the sequence in which they were transmitted. A Class 1 connection guarantees the availability of the full bandwidth. A port thus cannot send any other frames while a Class 1 connection is open.

Class 2 and Class 3, on the other hand, are packet-oriented services (datagram services): no dedicated connection is built up, instead the frames are individually routed through the Fibre Channel network. A port can thus maintain several connections at the same time. Several Class 2 and Class 3 connections can thus share the bandwidth. Class 2 uses end-to-end flow control and link flow control. In Class 2 the receiver acknowledges each received frame (acknowledgement, Figure 3.16). This acknowledge-ment is used both for end-to-end flow control and for the recognition of lost frames. A missing acknowledgement leads to the immediate recognition of transmission errors byFC-2, which are then immediately signalled to the higher protocol layers. The higherprotocol layers can thus initiate error correction measures straight away (Figure 3.18).Users of a Class 2 connection can demand the delivery of the frames in the correct order.Class 3 achieves less than Class 2: frames are not acknowledged (Figure 3.17). Thismeans that only link flow control takes place, not end-to-end flow control. In addition, the higher protocol layers must notice for themselves whether a frame has been lost. The loss of a frame is indicated to higher protocol layers by the fact that an expected sequence is not delivered because it has not yet been completely received by the FC-2 layer. A switch may dispose of Class 2 and Class 3 frames if its buffer is full. Due to greater time-outvalues in the higher protocol layers it can take much longer to recognize the loss of aframe than is the case in Class 2 (Figure 3.19).We have already stated that in practice only Class 2 and Class 3 are important. In practice the service classes are hardly ever explicitly configured, meaning that in current Fibre Channel SAN implementations the end devices themselves negotiate whether theycommunicate by Class 2 or Class 3. From a theoretical point of view the two service classes differ in that Class 3 sacrifices some of the communication reliability of Class 2in favour of a less complex protocol. Class 3 is currently the most frequently used service class. This may be because the current Fibre Channel SANs are still very small, so that

frames are very seldom lost or overtake each other. The linking of current Fibre Channel SAN islands to a large SAN could lead to Class 2 playing a greater role in future due toits faster error recognition.

 

 

 

 

 

KNOW MORE ABOUT RAID 2 and RAID 3

                         KNOW MORE ABOUT RAID 2 and RAID 3
When introducing the RAID levels we are sometimes asked: 'and what about RAID 2and RAID 3?'. The early work on RAID began at a time when disks were not yet veryreliable: bit errors were possible that could lead to a written 'one' being read as 'zero'or a written 'zero' being read as 'one'. In RAID 2 the Hamming code is used, so that redundant information is stored in addition to the actual data. This additional data permits the recognition of read errors and to some degree also makes it possible to correct them. Today, comparable functions are performed by the controller of each individual hard disk,which means that RAID 2 no longer has any practical significance.
Like RAID 4 or RAID 5, RAID 3 stores parity data. RAID 3 distributes the data of a block amongst all the disks of the RAID 3 system so that, in contrast to RAID 4 or RAID 5, all disks are involved in every read or write access. RAID 3 only permits the reading34 INTELLIGENT DISK SYSTEMS and writing of whole blocks, thus dispensing with the write penalty that occurs in RAID
4 and RAID 5. The writing of individual blocks of a parity group is thus not possible. In addition, in RAID 3 the rotation of the individual hard disks is synchronized so that the data of a block can truly be written simultaneously. RAID 3 was for a long time called the recommended RAID level for sequential write and read profiles such as data mining and video processing. Current hard disks come with a large cache of their own,which means that they can temporarily store the data of an entire track, and they have significantly higher rotation speeds than the hard disks of the past. As a result of theseinnovations, other RAID levels are now suitable for sequential load profiles, meaning that RAID 3 is becoming less and less important.2.5.6 A comparison of the RAID levelsThe various RAID levels raise the question of which RAID level should be used when.Table 2.1 compares the criteria of fault-tolerance, write performance, read formanceand space requirement for the individual RAID levels. The evaluation of the criteria canbe found in the discussion in the previous sections. CAUTION PLEASE: The comparison of the various RAID levels discussed in this section is only applicable to the theoretical basic forms of the RAID level in question.In practice, manufacturers of disk subsystems have design options in
• the selection of the internal physical hard disks;
• the I/O technique used for the communication within the disk subsystem;
• the use of several I/O channels;
• the realization of the RAID controller;
• the size of the cache; and
• the cache algorithms themselves.
The performance data of the specific disk subsystem must be considered very carefully
for each individual case. For example, in the previous chapter measures were discussed
Table 2.1 The table compares the theoretical basic forms of the various RAID levels. In
practice there are very marked differences in the quality of the implementation of RAID
controllers
RAID level Fault-tolerance Read performance Write performance Space requirement
RAID 0 none good very good minimal
RAID 1 high poor poor high
RAID 10 very high very good good high
RAID 4 high good very very poor low
RAID 5 high good very poor low2.6 CACHING: ACCELERATION OF HARD DISK ACCESS 35
that greatly reduce the write penalty of RAID 4 and RAID 5. Specific RAID controllers
may implement these measures, but they do not have to.
Subject to the above warning, RAID 0 is the choice for applications for which the
maximum write performance is more important than protection against the failure of a
disk. Examples are the storage of multimedia data for film and video production and the
recording of physical experiments in which the entire series of measurements has no value
if all measured values cannot be recorded. In this case it is more beneficial to record all
of the measured data on a RAID 0 array first and then copy it after the experiment, for
example on a RAID 5 array. In databases, RAID 0 is used as a fast store for segments in
which intermediate results for complex requests are to be temporarily stored. However, as
a rule hard disks tend to fail at the most inconvenient moment so database administrators
only use RAID 0 if it is absolutely necessary, even for temporary data.
With RAID 1, performance and capacity are limited because only two physical hard
disks are used. RAID 1 is therefore a good choice for small databases for which the
configuration of a virtual RAID 5 or RAID 10 disk would be too large. A further important
field of application for RAID 1 is in combination with RAID 0.
RAID 10 is used in situations where high write performance and high fault-tolerance
are called for. For a long time it was recommended that database log files be stored on
RAID 10. Databases record all changes in log files so this application has a high write
component. After a system crash the restarting of the database can only be guaranteed if all
log files are fully available. Manufacturers of storage systems disagree as to whether this
recommendation is still valid as there are now fast RAID 4 and RAID 5 implementations.
RAID 4 and RAID 5 save disk space at the expense of a poorer write performance.
For a long time the rule of thumb was to use RAID 5 where the ratio of read operations
to write operations is 70 : 30. At this point we wish to repeat that there are now storage
systems on the market with excellent write performance that store the data internally usingRAID 4 or RAID 5.
Buy Vmware Interview Questions & Storage Interview Questions for $150. 100+ Interview Questions with Answers.Get additional free bonus reference materials. You can download immediately even if its 1 AM. You will recieve download link immediately after payment completion.You can buy using credit card or paypal.
----------------------------------------- Get 100 Storage Interview Questions.
:
:
500+ Software Testing Interview Questions with Answers are also available plz email roger.smithson1@gmail.com if you are interested to buy them. 200 Storage Interview Questions word file @ $97

Vmware Interview Questions with Answers $100 Fast Download Immediately after payment.: Get 100 Technical Interview Questions with Answers for $100.
------------------------------------------ For $24 Get 100 Vmware Interview Questions only(No Answers)
Vmware Interview Questions - 100 Questions from people who attended Technical Interview related to Vmware virtualization jobs ($24 - Questions only) ------------------------------------------- Virtualization Video Training How to Get High Salary Jobs Software Testing Tutorials Storage Job Openings Interview Questions

 Subscribe To Blog Feed