#5137 WHO/TB/85.145 ENGLISH ONLY 1 July 1985 SURVEILLANCE OF TUBERCULOSIS BY MEANS OF TUBERCULIN SURVEYS by H.G. ten Dam, Scientist Tuberculosis and Respiratory Infections Unit WHO, Geneva, Switzerland INTRODUCTION Infection with tuberculosis is conveniently utilized as an index of the tuberculosis problem in epidemiological terms. The tuberculin test is a simple method of detecting infection, and, since infection is far more frequent than disease, fairly precise information can be obtained by examining a relatively small population sample. Long experience in tuberculin surveys has shown that the interpretation of the tuberculin test results may not be as easy as one would wish, for example if there is a high coverage with BCG vaccination, or a high prevalence of "non-specific" sensitivity. These problems will be discussed here; a further objective of the present paper is to provide guidance in the selection of an appropriate p0pulation sample. In the past, tuberculin surveys in developing countries were carried out mainly to obtain a rough idea of the magnitude of the problem. Whereas in this matter they may have largely met their purpose, they only rarely were sufficiently precise to provide appropriate base-line material for further studies. Often a sample was drawn that was only presumed to be representative for the population, and data were obtained for broad age groups. For surveillance the requirements are far more stringent. Surveillance implies that comparisons have to be made of the observed prevalences in selected age groups. Clearly, the age groups will have to be fairly narrow to exclude age-linked bias. When it comes to making comparisons, it will be necessary to determine which significance is to be attached to the observed differences. Therefore, a measure of the margin of error should be obtained with each estimate, which requires that the sampling should be carried out according to sound statistical principles. Some of the basic aspects of sampling are given in Appendix I. They will allow the reader to appreciate precision in sampling in relation to the size of the sample; the importance of representativeness and the advantages of cluster sampling; the proper utilization of stratification; and the analysis of results, in particular the comparison of data sets. Other aspects will be mentioned among the practical guidelines given below. Ce document iie constitue pas une publication. || ne doit faire I’objet d'aucun compte rendu ou résumé ni d'aucune citation ou traduction sans l'autorisation de I'Organisation mondiale de ia Same, Les opinions exprimées dans les articles signés n’engagent que ieurs auteurs, WHO/TB/85.145 page 2 THE TUBERCULIN TEST At the beginning of this century, von Pirquet(1) applied to tuberculosis an observation made in connexion with smallpox vaccination, i.e. that whereas primary vaccination caused no local skin reaction at all in the first two days, revaccination caused erythema within 24 hours. He demonstrated that if a tuberculous child was inoculated with a dilution of Old Tuberculin, a papule of 5 to 20 mm diameter appeared at the site of inoculation and disappeared within eight days. Apart from children already cachectic and those in the terminal stages of miliary tuberculosis and meningitis, children with clinical tuberculosis showed a positive reaction but in children with a different disease (and in whom tuberculosis had not been found on autopsy) the "allergy test" was negative. For many years the test was used in hospital patients with the aim of discriminating between the various forms of tuberculosis, or between a good and an infaust prognosis. (It had been remarked by von Pirquet that the most striking reactions were observed in bone tuberculosis and scrofula.) This application brought about the first quantitative method of tuberculin testing: by considering a reaction positive if its diameter exceeded a certain size (usually 5 mm) and administering successive tests with systematically increased doses, patients were classified according to the lowest degree of tuberculin dilution that elicited a positive reaction. A development connected with this was the introduction of the intradermal test, according to Mantoux, by which an accurately measured dose could be administered.* This quantitative method of testing, or "allergometry" as it was sometimes called, made it possible to determine fairly accurately the level of tuberculin sensitivity in patients, but it gradually appeared to be of little clinical value because the various levels of tuberculin sensitivity were not clearly correlated with different forms of tuberculosis, or with the severity of the tuberculous lesions. Hart,(3) investigating the sensitivity to tuberculin caused by tuberculosis infection, tested non-tuberculous hospital patients in parallel with tuberculosis patients with successive increased doses of Old Tuberculin and found that a dilution of 1/1000 was sufficiently strong to elicit a positive reaction in nearly all the tuberculosis patients. Moreover, in non-tuberculosis patients, additional tests with higher doses disclosed not more than a few per cent. additional reactors. ' ' Furcolow et al.(4) performed similar studies with PPD. Tuberculosis patients, their contacts and persons with a history of tuberculosis reacted to a dose of 0.0001 mg, but, Surprisingly, it was found that persons not reacting to this dose almost invariably reacted to much higher doses (up tp } mg). The hypothesis that such reactions were non-specific was. 5 . . . .confirmed by Palmer et al. when they feund that the two grades of tuberculin senSit1V1ty were distributed independently over the United States of America. *To simplify von Pirquet's test, several more or less ingenious variants have been proposed, such as various patch tests, the Heaf test, the tine test, etc., which aim at introducing into the skin a uniform, though unknown, dose of tuberculin and provide directly a negative or positive (sometimes graded) result. In practice these tests are satisfactory for screening, e.g. before BCG vaccination, but they are not indicated for measuring accurately the degree of tuberculin sensitivity. wHo!IBI85.145 page 3 In these studies by Palmer and his co-workers, the grade of tuberculin sensitivity in each person was not only determined by using successive increased doses, but the reactions to each dose were carefully measured and their appearance was described. From these and subsequent studies it became apparent that even if just one suitable dose of tuberculin was given, and the reactions were measured with consistent care, the distribution by reaction sizes of the tested population showed regular and meaningful shapes if presented in histograms, i.e. in a series of bars that have for their width different reaction sizes and for their height the percentage of all examined persons with that reaction size. Thus, the reaction sizes in tuberculosis patients were found to be distributed normally (in the statistical sense), whereas in a general population the distribution had a clearly bimodal shape (see Fig. l), with a right part (similar to the distribution in tuberculosis patients) representing the infected persons, and a left part (similar to a distribution observed after giving a placebo) representing the non-infected persons. The observed antimode of the distribution marks the point on the size scale (in Fig. l at 8—12 mm) that is best used as the criterion for distinguishing between infected and non-infected persons. FIG. 1. DISTRIBUTION OF TUBERCULIN REACTION SIZES IN TUBERCULOSIS PATIENTS (TOP) AND IN A GENERAL VILLAGE POPULATION (BELOW) 30 _ 2O _ 10 - 0 ' m m , n”... 3O — P e rc e n ta g e 20 ',_.- Tuberculin reaction (mm) In populations where low-grade ("non-specific“) tuberculin sensitivity is present thedistribution of reaction sizes assumes a completely different shape: bimodality may nolonger be observed (see Fig. 2). In such a population it remains possible to detect the FIG. 2. DISTRIBUTION OF TUBERCULIN REACTION SIZES IN UNVACCINATED CHILDREN, AGED 8-12 YEARS, IN A TROPICAL COUNTRY P er ce nt ag e Tuberculin reaction (mm) WHO/TB/85.145 page 4 non-infected persons, i.e. to distinguish between those who had just a traumatic reaction to the test and those who had low-grade sensitivity, by giving a second test with a higher dose to all persons with a low or intermediate reaction to the first test. On the other side it is far more difficult to discriminate between those with non-specific and those with specific sensitivity. Logically this would require a second test with a product that elicits a reaction to only one of these kinds of sensitivity. So far such a product has not been found and only some ways of approximation are aVailable for making the desired distinction. INTERPRETATION OF THE TEST For several reasons the quantitative tuberculin test develo ed by Palmer and his co-workers of the WHO Tuberculosis Research Office(6’ 7’ é, 9’ 10 , has become an important epidemiological tool. One of the reasons is that tuberculosis infection is invariably far more prevalent than disease so that it (or any transformation of data from tuberculin testing that may be considered a better characteristic of the tuberculosis problem) can be determined more easily. A further reason is that it can be determined in easily accessible age groups (children) in which it would be very difficult to estimate the prevalence of disease; clinical tuberculosis of the adult type is rare in children and since disease causes absenteeism, school children would not constitute a representative group, even in countries where school attendance is complete. But perhaps the most important reason is that data on the prevalence of disease collected by different investigators are hardly ever comparable, owing to the large variations in diagnostic criteria and techniques, whereas the tuberculin test provides data that are comparable in time, and from place to place, even if different tuberculin products are used and different persons read the reactions; since the test results are interpreted according to the distribution observed, there is no strict need for standard procedures, supplies and equipment. In general, the interpretation of the results of tuberculin testing poses no problem but, as mentioned above, in countries where non-speciic sensitivity is highly prevalent the interpretation may be difficult. The fact that M. bovis and BCG induced the same kind of sensitivity as M. tuberculosis had always been accepted as a shortcoming that did not diminish the epidemiological value of the test since infection caused by M. bovis for practical purposes was considered similar to that by M. tuberculosis and it was conSLdered easy to exclude BOG-vaccinated persons because vaccination leaves a characteristic scar. Sensitization by yet other micro-organisms, however, seriously reduces the predictive value of the test since the interpretation of the observed distribution of reaction sizes becomes largely speculative. In a population where only specific sensitivity to M. tuberculosis or M. bovis is prevalent, the antimode of the distribution can be taken confidently as the test criterion for distinguishing between the infected and the non-infected. This antimode (if observed at all), however, has no epidemiological significance if non-specific sensitivity is present, i.e. if the observed distribution of reactions is the combination of three (or more) overlapping distributions in epidemiologically different groups. In fact, it was even found that an antimode may be provoked by the in vivo effect of a stabilizing agent in the tuberculin preparation (Tween 80) which tends to reduce small reactions and increase large ones irrespective of the cause of the sensitivity. Obviously in such cases it is not possible to tell whether the sensitivity observed in a certain individual is specific for tuberculosis infection or not. For epidemiological purposes, however, it is sufficient that the ro ortion of infected persons can be estimated. One approach is to assume that all non-specific senSLtivity produces reactions of intermediate size, and thus that the right-hand tail of the distribution represents infected persons only. Since in infected persons the distribution of reaction sizes is approximately normal, the left-hand tail of the distribution of reaction sizes in the infected persons can be plotted if the mean can be estimated, e.g. from the reactions observed in patients. (Patients react to the tuberculin test in the same way, whether non-specific sensitivity is present in a c0untry or not.) WHO/TB/85.145 page 5 Another approach is to try and distinguish specific and non-specific sensitivity by applying a second test. Unfortunately, efforts to produce a tuberculin thaf is more specific for infection with M. tuberculosis than the PPD prepared by Seibert in 1934 11 have failed so far. In fact, most workers seem to have been so intrigued by the non-specific sensitivity that they concentrated on investigating the nature of this rather than on distinguishing it from specific sensitivity. A number of PPD products (mycobacterins) have been prepared from mycobacteria other than the tubercle bacillus. In a suitable dose, these products, such as the ones prepared from M. avium and M. intracellulare, elicit in certain persons a larger reaction than the tuberculin test, but in others a smaller reaction. The variations are so large and systematic that they cannot be attributed to experimental error, as would be observed if two identical tuberculin tests were given simultaneously. The assumption can therefore be made that if the tuberculin produces the larger reaction, the person is infected by virulent tubercle bacilli, but if the (simultaneously administered) sensitin produces the larger reaction, the person's sensitivity is non-specific. Obviously, this dual testing must be applied and interpreted with great care. The mycobacterin may indeed be more specific than tuberculin for the mycobacterium from which it was prepared, but this mycobacterium may not be the (only) cause of non-specific sensitivity in the population. The dosages of the tuberculin and the mycobacterin must be balanced in a way that is suitable in the particular population (or, preferably, the interpretation must be made according to the observed correlation). The result of the testing can be interpreted with a reasonable degree of confidence when the difference between the two reactions is large but becomes problematic when the reactions are about the same. (The assumption that in that case both specific and non-specific sensitivity are present means systematically classifying all persons in this group as infectedl) It will be clear that if the difference between the two reactions is to be large, the reaction to tuberculin has to be either small or large° It is precisely such reactions that can be interpreted fairly confidently also in the absence of a second test result. Thus, the additional information obtained from the second test is very small. In fact, the only definitive value of the second test is that it helps to distinguish between reactions that are the results of some form of sensitivity and those that are simply trauma caused by the injection. In this respect the test with a mycobacterin is generally less powerful than a second tuberculin test with a higher dose, but it has the advantage that it can be administered simultaneously. More recently a method has been proposed(12) by which the selected population is given a tuberculin test and is (re-)vaccinated simultaneously with BCG. A further tuberculin test is given 8-12 weeks later. The result of two tests are entered in a correlation table. Interpretation is as follows: Persons infected with tuberculosis will, on the average, have the same reaction to both tests because repeated testing and BCG vaccination do not "boost" specific sensitivity. A strong dose of B00 in addition to the first test however, will cause an increase in sensitivity when the first reaction is non-specific, especially if the skin sensitivity had waned which is common in sensitization with atypical mycobacteria and after BCG vaccination. Inspection of the correlation table will make it possible to identify the persons with tuberculosis infection or at least make an estimate of their proportion in the study population. This method seems particularly indicated in populations where vaccination is given at birth and re-vaccination at school age. It should be noted that all test results in some way are interpreted by making use of the actual findings in the population tested. Using a "standard" criterion, i.e. one based on observations from elsewhere, no doubt may seem simpler but is far less accurate. It has no place in surveillance where results obtained at different times have to be compared. PRODUCTS AND DOSAGE Tuberculins may differ qualitatively; for instance, two products that give similar reactions in infected guinea-pigs may give different reactions in BOG-vaccinated children. In addition, they may differ in physical properties, such as in the adsorption to glass. These differences make it hard to express their strength in equivalence of the adopted international units (separate ones were established for OT and PPD). These problems, ho ever, are of little practical importance; for epidemiological studies it is sufficient to obtain a tuberculin in a strength that has been shown to be suitable for use in human beings and to apply a quantitative method as described above; also, for the evaluation of ECG WHO/TB/85.145 page 6 vaccines the use of carefully calibrated products (or always the same product) alone does not solve the problem because other, uncontrolled, variations in experimental error are often of greater importance. For this reason the comparability of studies on BCG vaccines is best ensured by including systematically the same BCG reference preparation in the investigations. This of course does not imply that all tuberculins are equally suitable for all purposes. Qualitative differences, such as in the specificity for detecting tuberculosis infection, would be of great importance. Unfortunately, in this respect there is no outstanding product among the currently available preparations. The rather crude OT preparations, which still include many components of the culture medium, are generally considered less suitable than PPD. Several workers who originally utilized OT were tempted to continue working with it because they wanted to enSure comparability in their material. With the quantitative method of testing, this argument no longer applies and OT preparations should be considered obsolete. A special large batch of PPD (RT23) was prepared by agreement with UNICEF and WHO by the Statens Seruminstitut, Copenhagen, Denmark. This WHO—owned product is usually supplied in isotonic dilutions ready for use (in rubberstoppered vials) which have been stabilized with Tween 80 to prevent adsorption of the tuberculin to the wall of the glass containers. This preparation was assayed against the International Standard for PPD and extensively compared with other products. The usual human dose is commonly, but incorrectly, denoted as "2 TU". For other tuberculin preparations which may not contain a stabilizing agent, the strength is usually expressed in equivalents of International Units, and a suitable human dose is 3 to 10 I U. There is an International Standard for avian "tuberculin", and the strength of such preparations may be expressed in International Units. The correctly specified usual human dose is 5 I U; for the preparation stabilized with Tween 80 the dose has been (incorrectly) denoted as "2 TU". There are no International Standards for the various other mycobacterins; the ones mentioned are "calibrated" against RT-23 and PPD-S respectively on a weight-to-weight basis, and presented in dilutions ready for use. Sometimes their strength is also indicated in "Units" or (incorrectly) as "TU". The strength provided may not be the optimum strength for dual testing. Mycobacterins prepared from mycobacteria other than the tubercle bacillus prepared in a way similar to RT-23 are available commercially from the Statens Seruminstitut, Copenhagen. The products available through WHO, and their usual human dose (i.e. per 0.1 ml, for intradermal injection) are: Product Dose Presentation RT 23 (with Tween) 2 TU dilutions ready for use; stock solution of lg/L; PPD powder PPD-S (for epidemiological 3-10 I U freeze-dried in ampoules studies only) (+ ampoules of diluent) Tuberculin dilutions keep well at ambient temperatures. They usually have an expiry period of half a year, but have been reported to deteriorate in ultraviolet light, so that extended exposure to sunlight should be avoided. EQUIPMENT AND STERILIZATION Syringes and needles for tuberculin testing are kept separate from those used for other purposes such as BCG vaccination. Leak-proof graduated syringes such as "Omega microstat" are used, with platinum or steel, 25 or 26 gauge, 10 mm needles.* *Platinum needles can be flamed to red-hot without damage and remain sharp. Steel needles quickly become corroded when used in this way, which makes injection painful, but they are far cheaper and thus could be replaced more frequently, for instance daily. WHO/TB/85.145 page 7 Before each session syringes and needles are disassembled and are washed and brushed with water and then sterilized in an autoclave. If this is not possible, they are boiled for at least 10 minutes in distilled or demineralized water (as used for car batteries) to avoid calcium deposits. During a single session a syringe and needle may be used for many injections, provided the needle remains on the syringe. If it becomes detached, it has to be sterilized as mentioned above. The tip of the syringe is flamed; the sterilized needle is taken up with forceps, the hub is heated and the needle is mounted on the syringe in such a way that the bevel is at the side of the graduated scale. Heating the hub of the needle (rather than just flaming it) will cause a slight expansion and consequently a tight fit on the syringe when the needle has cooled down. The whole needle is flamed before the syringe is filled from the vial or ampoule containing the tuberculin. Air remaining in the syringe after filling is expelled. Before each injection the tip of the needle is flamed to dull red; 3 small amount of tuberculin is expelled to cool the needle and to make sure it is not obstructed. Normally a clean skin is not sterilized before injection. The reactions are measured with a small transparent ruler graduated in mm. A metal box to keep the syringes and needles in, a small spirit lamp for flaming and suitable forceps to mount the needle on the syringe are also needed. The syringe box and the forceps should be cleaned and sterilized together with the syringes and needles. ADMINISTRATION OF THE TEST The tuberculin test is usually given on the dorsal side of the forearm, but the volar side may also be used. The test should not be given at a site previously used for tuberculin testing. It is advisable to avoid such a site for several years. If two tests have to be given simultaneously, one is given on each arm. The different tests are preferably not given systematically to either the left or right arm to avoid bias in reading but allocated according to some random procedure, which, of course, should be recorded accurately for each person. The tip of the needle (with the bevel upwards) is inserted, lengthwise to the arm, superficially into the skin which is slightly stretched in the opposite direction. The syringe is held by the barrel only. When the needle has been inserted satisfactorily, the position of the piston ring is observed, the syringe is held firmly in place, and the piston is actuated gently until the piston ring has reached a point 0.1 ml ahead of the original position. Then the needle is withdrawn, the syringe being held parallel to the arm. The injection should be given slowly so that tissue damage is avoided as far as possible. A well-given intradermal injection of 0.1 ml will produce a raised but flat anaemic weal with an "orange peel" aspect and a diameter of some 7 mm. If the injection has been given too deeply, the weal will be smaller, dome-shaped and less anaemic. This will make the reaction more difficult to read. Whereas it is generally true that a correctly given intradermal injection of 0.1 m1 gives a weal of about 7 mm, the converse, i.e. that a weal of 7 mm indicates that 0.1 m1 has been injected correctly, is obviously not true. The old custom of estimating the dose from the weal produced was therefore very inaccurate and could only be excused by the fact that the previously used types of syringe leaked so much that the reading on the scale was perhaps even more inaccurate. When the test has been administered, this fact is recorded, e.g. by entering the date on the record form. The intradermal administration of the tuberculin test usually poses no problem but even the most experienced tester (especially when testing small children) will experience from time to time a technical incident resulting in an inadequate administration of the test dose. Usually in such cases no attempts at correction are made; only if the incident happens before any tuberculin has been injected the test may be given again at a different site. Whereas an occasional mishap is unavoidable, it will usually be of little consequence provided the mishap is clearly recorded so that the test result will not be interpreted as though the test were given correctly. Testing should be as uniform as possible. If two or more testers work in the same programme, they should be allocated judiciously to the population to be tested (see underReading below). WHO/TB/SS. 145 page 8 To avoid bias in reading, the reader should not be informed of the possible sensitivity of the person examined; for instance, he should not know the vaccination status of the person examined, or know the result of a previous tuberculin test. Thus, he should read the tuberculin reaction before even looking (or enquiring about) the presence of a BCG lesion. If two tests have been given to each person (e.g. a tuberculin and a sensitin) in different arms, the reader will first examine the reaction on one arm in a number of persons and then the other in these people. The reader should not have access to records that may give information on previous test results. In general, it is highly advisable that the records are kept by a registration clerk, to whom the reader dictates his findings. ORGANIZATION OF A SURVEY A central office is required where the work is planned, prepared, and analysed. Field teams are needed to carry out the actual examinations. The work in the central office will require one medical officer who should be well acquainted with field work in general as well as with sampling. The field team should consist of at least one technician (leader) experienced in field work, one registration clerk, and a driver if the survey is in a rural population. The leaders of the field teams and the registration clerks should receive, in the central office, adequate training in the procedures to be followed, especially in the final stages of selecting the samples, in recording, and in reporting. Special training in tuberculin testing, and in reading the reactions, will be required, and the skill in these techniques should be assessed occasionally, e.g. by double readings (same reader) and by correlating the readings of one reader with those of another reader. READING THE REACTIONS A reaction to tuberculin has different attributes such as induration, surrounding oedema, erythema, density ("firmness"), presence of bullae or vesicles or of lymphangitis and fever. All these attributes reflect the degree of sensitivity and depend on the dose of tuberculin given. Moreover they are all correlated, although for different products the correlations may vary. Thus for routine testing the measurement of the tuberculin reaction can be restricted to one of the attributes. The most practical ones to be considered are induration or erythema, and it is probably unimportant which of the two is used. It has been suggested that it is easier to measure the induration after PPD and the erythema after 0T. Since OT has been superseded and since it is difficult to measure erythema on a very dark skin, the measurement of induration is now generally preferred. The test is read three days after it has been given. The test site is carefully palpated and if induration is present its diameter is measured with a tranSparent ruler in a uniform way, e.g. transverse relative to the arm, and recorded in millimetres. If there is no induration, "0” is recorded. The readings should be made by trained observers so that the readings are both accurate and consistent. If two or more readers work in the same programme, small, systematic differences in the readings may occur even if the readers are equally accurate and consistent. For this reason the differences between readers should be measured. This is achieved by judiciously allocating the tested population to the different readers. For instance, if post-vaccination reactions have to be read by two readers in school-children vaccinated with two different batches of vaccine, it would be appropriate to allocate by some random procedure half the children vaccinated with either batch to one reader and the other half to the other reader. If this were done, not only could the readers‘ work be compared but also the effect of possible differences in reading on the comparison of batches would be minimized. Planning of surveillance should be based on clear objectives. In general it will be advisable to start with a base-line survey that covers the entire country, even if the national tuberculosis programme does not yet cover all areas. It may however be possible that certain areas appear completely inaccessible or have an exceptionally low population density (deserts, flooded territories) so that they would have to be excluded if it so happened that they were selected in the sample. Such areas must be excluded a priori; alternatively a list is made of the areas to be included in the base-line survey. TABLE I. SAMPLE SIZE REQUIRED TO SHOW A SIGNIFICANT DIFFERENCE (5% LEVEL; P° = 0.80) BETWEEN INFECTION PREVALENCES AFTER 5 YEARS, ACCORDING TO INITIAL ANNUAL RISK OF INFECTION, EXPECTED ANNUAL DECREASE IN RISK; AND AGE OF STUDY POPULATION Initial risk Annual decrease Children aged 6 yrs Children aged 10 yrs 2% 85 200 145 200a 4% 22 400 38 70014 6% 10 000 17 4008% 5 000 10 700 22 42 200 76 40042 11 200 20 2002% 6% 5 200 9 4008% 3 000 5 500 22 29 300 53 100 4% 7 600 13 8003% 6% 3 500 6 4008% 2 100 3 800 Children of school-entrance age would appear suitable for surveillance, also forpractical reasons. SAMPLING Selection of the sample is initiated in the central office. Both demographic data and,if possible, previous data on the prevalence of tuberculosis should be considered. Effortsshould be made to obtain such information in detail, not only from official census reportsbut also from recent surveys and programmes. If the survey is actually carried out in schools, an adequate method would be to obtain alist of all schools with the numbers of pupils in the selected age group and to randomly drawa sample that would give the required study population size. In many, eSpecially large,countries this information may not be easily obtainable and it may be preferable to base thesampling in the first place on general population figures. This obviously also applies ifthe proportion of children attending school is low, and especially if it differs in variousparts of the country. Figures for the selected age group, if available, are used inpreference to figures for all ages. WHO/TB/BS . 145 page 10 The demographic information is used in the first place to exclude areas which are inaccessible to the field team. These areas may be marked off on a map. Then the epidemiological information is examined to determine whether it is profitable to divide the remaining population in different strata. In this connection it should be borne in mind that it is a disadvantage to include strata, that would turn out not to differ much from the national average. The number of strata therefore, should always be small. Also the population in a particular stratum should be at least a fair proportion of the total population; in a national survey it is convenient to attribute the sample population proportionate to the population in the different strata. Experience has taught that it may be profitable to divide the population in an urban and a rural stratum, the definition of what is urban or what is rural has to be made in the central office, if possible based on epidemiological considerations. Any stratum should include at least two clusters. Once the strata have been defined the sample to be allocated to each stratum is determined, usually by taking it proportional to the population. The next task is to define for each stratum the areas in which the sample is to be taken. A practical method is to assign to each stratum a number of clusters that would yield the planned sample size. If, for instance, it has been decided to obtain a national sample of 20 000 children and 60% of the pOpulation in the country is rural, 12 000 children are to be found in the rural stratum. If a cluster of 200 children aged 6 years might be found in a population of some 10 000, 60 areas of some 10 000 population would have to be selected. This may be done in several steps. To obtain a first gross allocation of clusters, each stratum is divided into administrative areas (say districts) for which population estimates are available. The clusters are allocated according to the population in the different districts. This is done by accumulating the population and drawing from a list of random numbers as many numbers as there are clusters (60), between 0 and the population total. For instance if in a stratum of 12 million population there are 20 districts, the population figures are accumulated as in column (3) below: (1) (2) (3) (4) District Population Accumulated total Clusters selected 1 450 000 450 000 2 530 000 980 000 3 620 000 l 600 000 4 480 000 2 080 000 l 20 570 000 12 000 000 If 60 clusters are to be drawn, 60 random numbers between 0 and 12 000 000 are drawn from a list of random numbers. For instance in the Statistical Tables by Fisher and Yates (Sixth Edition), on page 138, left hand four columns going down, we find: Random number No. of cluster 28 896 587 30 294 365 95 746 260 01 855 496 l 10 914 696 2 Of these figures the fourth and the fifth are between 0 and 12 000 000 and thus accepted. The fourth figure designates District 4 in the example given above, since it is more than 1 600 000, but less than 2 080 000. WHO/TB/85.145 page ll It should be noted that in this random way of selection several districts will be allotted more than one cluster, whereas some may be allotted none. It may therefore seem advantageous to draw a single random number (in this case below 200 000) and to proceed from there in intervals of 200 000. This systematic way of sampling has only a slight administrative advantage, but will result in a larger (not, as may have been hoped, a smaller) confidence interval of the final estimate because variation between clusters will be larger than with random allocation. Once the districts have been selected, the procedure is repeated to locate the cluster(s) within each district. The districts are divided into 5-10 areas for which the population can be estimated. The populations are again accumulated and accordingly the cluster(s) are allocated randomly to the areas. Finally these areas are divided in blocks of the cluster size, and the clusters are selected randomly. This final selection may have tobe done by the field teams. A method which is not recommended, but still frequently used, is to find by some random procedure a certain house—in a selected area and then proceed to the nearest house, etc., until the required cluster size has been obtained. The main disadvantage of this procedureis that the cluster size is no longer self—balanced, i.e. it is not proportionate to the number of people actually living in the selected area. When the area is so large that it would be cumbersome to divide it entirely into blocks, it may first be divided into a convenient number of well-defined sectors, for instance of some 50 000 inhabitants; the population totals for the sectors are accumulated, and one sector is selected randomly for each cluster allocated. The sector thus chosen is thendivided into blocks and one block is selected. This procedure may also be followed when thearea has been allotted several clusters. In this case the selection of a sector does notexclude it from further selection. All children in the eligible age group living in the block belong to the cluster andtherefore have to be included in the study. If the children are examined at school it shouldbe verified that they indeed live in the selected block. When school attendance is low,home~visits may be needed to supplement the coverage of the population. If it has been decided to include only schoolchildren in the survey the allocation maybe carried out as far as the areas. For the areas selected the schools are listed and aproportion of them is selected that would provide the estimated cluster size. FIELD WORK The central office should facilitate the work of the field teams as much as possible.Important gains in time and accuracy are made if the field visits are properly prepared. Thefield teams must have received clear instructions where to go (an efficient itinerary throughall clusters should be prepared). ApprOpriate equipment and supplies, adequate demographicdata (maps), and if possible names of local contact persons should be provided. The visit ofthe field teams should be announced beforehand, and the purpose should be duly explained tothe local authorities. The team should be provided with an explicit letter of introduction. On arrival at the place of the cluster sample, the field team contacts the localadministration, the village head, the school master, etc. The purpose of the visit isexplained, any information required is collected, and if possible the cooperation of a personwho is familiar with the local circumstances is ensured. If a house-to-house survey is carried out, a map is prepared (if a map cannot beobtained locally) then all households or dwellings included in the sample are identified andvisited. Their exact location (household No.) is indicated on the map. (It may beconvenient to paint the household number on the house as well). WHO/TB/85.145 page 12 In each household the team leader explains the purpose of the visit and asks the head or a responsible person to give the name, age and sex of all persons in the household in the eligible age group, to start with the youngest. The data for persons in the eligible age group are recorded by the registration clerk, including for persons who are temporarily absent, on a household registration form (Fig. l) for every household. The form is self-explanatory. The forms are numbered consecutively for each cluster. The cluster number will have been provided by the central office. FIG. 1 HOUSEHOLD REGISTRATION FORM Date: Cluster No. Household No: Name: Address: Tester/reader: Persons examined: Registration clerk: Tuberculin test BCG Scar Size (mm) IndurationFirst Name Age Sex (mm)Given (date) Read (date) l. If the survey is carried out in schools only, registration and recordings are made on lists giving the name of the school and the designation of the class but otherwise similar to the household registration form. Registration and administration of the tuberculin test are carried out on the same visit. When all the persons in the age group eligible have been registered, they are given the tuberulin test. The date is filled in to show that the test has been given. When the tests have been given, all persons registered are examined for the presence of a BCG scar or lesion. If no BCG scar is found, a dash is put in the box for "Size". If there is induration or an ulcer (from recent vaccination) no measurement is recorded under "Size”, but "induration" or "ulcer" is filled in. If more than one lesion is visible, the diameter and appearance of the largest lesion is recorded. If a person is temporarily absent, no recording is made. If in the household there are no persons eligible for examination, "none eligible" is noted after "Persons examined". If the team cannot contact the household, the reason for this is also noted after "Persons examined", for instance "household absent" or "refused". The team will try to examine those who are temporarily absent at another time or another place (e.g. at school, with neighbours). The tuberculin test is read three days after it has been given, and the size of the induration is recorded. When the test cannot be given or read, the reason for this (e.g. absent, refused) is given. All information (including the name of the tester/reader) is recorded by the registration clerk who always keeps the cards. The tester/reader dictates his findings without referring to the card. WHO/TB/85.145 page 13 When the work in a cluster is finished, the team verifies the records. No corrections are made on the household registration forms, but any supplementary indications are noted under "Remarks" on the Cluster Report (see Fig. 2). FIG. 2 CLUSTER REPORT Date: Cluster No: District: Town or Sector: No. of households visited: Selected but not visited: No. of persons registered; No. tested: ' No. read: Examiner: Registration clerk: Remarks: The cluster report contains information that is common to all the households (District, town, village, sector) and a brief summary of the performance of the field team. Locally obtained or prepared information, such as population figures, maps, itineraries, etc., are attached to the cluster report, as is the set of completed household registration forms. Any additional observations may be recorded under "Remarks". The completed cluster report and attachments are submitted to the central office as soon as possible. Final work in the central office As soon as the Cluster Report is received it is examined for adequacy of coverage by the field team and of the sampling procedure. Inadequate coverage of the cluster may call for a repeat visit; flaws in the sampling procedure demand immediate corrective action. Due attention should be paid to the "Remarks" on the Cluster Report, which may indicate that different arrangements are to be made. Locally obtained maps, population figures, etc., are filed for future reference. For each cluster the data are analysed separately for a start. Histograms are drawn of the tuberculin reactions measured, separately for those with and without BCG vaccination. According to the distributions observed the proportions of persons infected for the different ages, are estimated. The infection prevalences thus obtained are utilized to obtain the required estimates as indicated in the Appendix, and a final report is prepared. WHO/TB/BSJAS page 14 9. 10. REFERENCES von Pirquet, C. (1907) Tuberkulindiagnose durch cutane Impfung, Berliner klin. Wschr., fig, 644 von Pirquet, C. (1907) Die diagnostische Wert der kutanen Tuberkulinreaction bei der Tuberkulose des Kindesalters auf Grund von 100 Sektionen, Wiener klin. WSchr., 29, 1128 ' Hart, P. D'Arcy (1932) The value of tuberculin tests in man, with special reference to the intracutaneous test, British Medical Research Council Special Report Series, No. 164 Furcolow, M.L. et a1. (1941) Quantitative studies of the tuberculin reaction. I. Titration of tuberculin sensitivity and its relation to tuberculous infection, Public Health Reports, 52, 1082 Palmer, C.E. et a1. (1950) Studies of pulmonary findings and antigen sensitivity among student nurses. VI. Geographic differences in sensitivity to tuberculin as evidence of nonspecific allergy, Public Health Reports, 22, 1111 Palmer, C.E. & Strange Petersen, 0. (1950) Studies of pulmonary findings and antigen sensitivity among student nurses. V. Doubtful reactions to tuberculin and to histoplasmin. Public Health Reports, 25, 1 Guld, J. (1953) Quantitative aspects of the intradermal tuberculin test in humans. 1. The dose-response function within the range 1-10 tuberculin units, determined by duplicate tests. Acta tuberc. scandinav., Zé, 222 Guld, J. (1954) Quantitative aspects of the intradermal tuberculin test in humans. II. The relative importance of accurate injection technique. Acta tuberc. scandinav., 29, 16 WHO Tuberculosis Research Office (1955) The S TU versus the 10 TU intradermal tuberculin test. Bull. w1d Hlth 0rg., 13, 169 Guld, J. (1957) Interpretation of tuberculin reactions in populations with a high proportion of ECG-vaccinated persons. Bull. Wld Hlth Org. 11, 225 11. Long, E.R. et a1 (1935) A standardised tuberculin (purified protein derivative) for uniformity in diagnosis and epidemiology. Tubercle, 19, 304 12. ten Dam, H. G. & Hitze, K. L. (1980) Determining the prevalence of tuberculosis infection in populations with non-specific tuberculin sensitivity. Bull. w1d Hlth 0rg., g, 475 Precision An estimate of the prevalence of infection has little value if one cannot state how good an estimate it is. Thus an estimate of, for instance, "23%", provides little information and even may be misleading, as it suggests precision. It is therefore necessary to provide the interval in which the true value is likely to lie. This interval depends on the data collected, but also on how confident one likes to be that the true value indeed lies in it. It is the custom to require at least a confidence level of 95%, i.e. the interval must be so wide that, on the average, in 95% of cases the true value is expected to lie in it. This interval is called the "95% confidence interval", and its limits the "95% confidence limits". Arithmetically the 95% confidence interval (1) depends on the number of persons included in the sample (n) and on the percentages of persons with and without the characteristic in the sample (p and lOO-p), according to the formula: 1 = p 1 1.96 / 2 (100-2) 1'1 If in a sample of 600 persons there are 120 persons classified as infected, 80 x 20= ° .____.__ v = o 0 °I 204 : 1.3?/ 600 m 404 i 3.24 and thus the higher and lower 95% confidence limits are 23.2% and 16.8% respectively. It should be noted that the size of the population from which the sample is drawn does not enter into the arithmetic, and therefore does not influence the estimate. This is because it has been assumed that the sample is drawn from an infinite population, or with "replacement", i.e. that each person selected can be selected again. It is seen that the width of the interval depends on the square root of the number of persons in the sample, so that a sample 4 times as large (if the same proportions had been found) would be given a confidence interval half as wide. Sample size In practice the question of precision also presents itself in a reverse form. Before a sample is selected one may like to know how many persons should be included, i.e. the minimum sample size has to be determined for a pre-determined confidence interval. In surveillance the question is slightly complicated by the fact that it will be necessary to compare the prevalences obtained in different surveys. The samples should be at least so large that the difference between the prevalences would be statistically significant. For a 90% chance of obtaining a significant difference at the 5% level, the study population in each survey (n) should be; 2 8.6 [p1 (lOO—pl) + p2 (100—p2)] 2 n (Pl—P2) where p1 and p2 are the prevalences in the two surveys. Obviously these values will have to be estimated. The formula applies to simple random sampling. For cluster sampling larger populations would be required. It is not possible to provide an estimate for this, since various factors would have to be considered. The best procedure is to select an ample WHO/TB/85.145 page 16 Appendix 1 population (No. of clusters) and review the situation when a good proportion has been covered in the first survey. Representativeness Ensuring that the sample drawn is representative demands the greatest precautions to exclude bias. The simplest way is to draw the sample at random. This means that (a) every person in the population that the sample is to represent should have an equal chance of being included, and (b) his selection should be independent of that of any other person. To select each individual randomly, for instance, by listing the whole p0pulation (or making use of existing lists) and drawing a sample of the required size according to a list of random numbers, is generally prohibitive in practice. The optimal solution is very often to make up a sample out of several randomly selected groups. Rather than a single person, a "cluster" of persons is the sampling unit. In this way the cost of selecting and examining each person in the sample may be greatly reduced. Within a cluster, variation would be less than in a random sample. The inclusion of several randomly selected clusters, however, makes it possible to estimate the variation between them, and thus, as in a random sample, to meaSure the margin of error. Cluster sampling In cluster sampling the 95% confidence interval of the estimate is: 2Si (pi p) k - 1) in the entire sample, pi the percentage found in cluster i, and k the number of clusters. , when p is the percentage found An important variation between clusters may be observed if the distribution of infection in the population is "patchy", e.g. when it differs widely from region to region. This heterogeneity will be reflected in the confidence interval: cluster sampling would give a larger confidence interval than simple random sampling, since the variation between clusters contributes to the standard error of the estimate. Stratification Greater precision may be obtained if it is possible to divide the population in some distinct strata in which the variation is small relative to that in the whole population. A number of clusters is then allocated to each of the strata. Examples of strata used in country surveys are different geographical areas, urban and rural populations, etc. It is interesting to note that because stratification ensures a higher representativeness than can be expected from random selection of clusters, it causes the variation between clusters to be larger. A gain in precision is made, however, because the estimate that can be made for each stratum will have a smaller standard error than that for the same number of clusters randomly selected in the country. When the different stratum estimates are pooled to obtain the estimate for the whole country, the differences between stratum estimates do not contribute to the standard error of the estimate. It will be clear that stratification must be carried out upon pre-existing information (i.e. not on the basis of the actual results obtained). The 95% confidence interval of the estimate will be: 2 I - p i 1.96 fig E; (psi ps) k(k-s) p is the percentage found in the entire sample Psi is the percentage found in cluster 1 of stratum 3 ps is the percentage found in stratum s k is the total number of clusters, and s is the number of strata , when WHO/TB/85.145 page 17 Appendix I Since the calculation of the confidence limits is based on the measurement of the variation in each stratum, it will be clear that at least 2 clusters should be selected in each stratum. The greater the homogeneity within strata and the heterogeneity between strata, the more efficient the stratification will be. It is only indicated to divide the population in strata if these conditions are likely to be found. The reader may verify that since the number of strata (5) occurs in the denominator of the formula, any unjustified stratification will result in an increased confidence interval. Usually the number of clusters allocated to each stratum is taken proportionally to the population in the various strata, since it avoids having to use weighted means when analysing the results. Number and size of clusters The optimum number of clusters depends largely on the homogeneity of the infection rate in the population. When deciding on the number of clusters, however, a first estimate can be made by determining the total sample size as for simple random sampling. If the distribution of infection in the population is expected to be very ”patchy", a few large clusters will give much less information than a large number of small ones; it therefore should be attempted to make the number of clusters as large as possible. In practice the daily output of the field team will be an important factor in determining the cluster size. If the work in one cluster will take one or more days, for instance, it will generally be desirable to finish work in one cluster at the end of the working day. Another important consideration is that of the proportions of the time spent on preparation (and transport) and on the actual examinations. Comparing prevalences In surveillance it will eventually be required to compare the estimated prevalences: This is done by determining the confidence interval of the difference. I 2 2 W1 + WZ, when p1 and p2 are the two prevalences observed, and W1 and w2 the widths of their 95% confidence intervals. N lr— I I = p1 ” p2 1 Thus, if in one year the prevalence in a certain age group is found to be 10%, with a confidence interval of 6% to 14%, and some years later it is found to be 4%, with a confidence interval 0f 1% to 7%, the 95% confidence interval of the difference between the prevalences is I = 10 - 4 i l J/SZ + 62 = 6 i 5, or +11% to + 1% 2 It may be concluded that there is a 95% chance that the second prevalence is 1% to 11% less than the first one. More roughly it may be said that the second prevalence is "significantly" less than the first (at the 95% level). Similarly an interval including only negative values would have indicated that the first prevalence had been significantly lower than the second. A particular situation arises if the confidence interval includes the value 0, i.e. that it includes both a positive and a negative value. If this occurs the difference is not significant, at least not at the 95% level. It must be pointed out that if values are found not to differ significantly, this does not mean that they are "about the same". Such a conclusion would be justified only if both confidence limits were close to O, in which case it is immaterial whether one or both are positive or negative.
Organisation mondiale de la santé (OMS) · Technical Documents
Surveillance of tuberculosis by means of tuberculosis surveys
Voir le document original
Le texte intégral est hébergé par l’organisation qui le publie. lawenc.com indexe les métadonnées et renvoie vers la source officielle.
Texte intégral
Informations clés
Organisation
Organisation mondiale de la santé (OMS)
Type de document
Technical Documents
Source
Organisation mondiale de la santé