Tuesday, September 8, 2026

SDTMIG-4.0 Draft - The new LC (Lab Test Results in Conventional Units) domain

 

CDISC just published a new set of draft SDTM domains for the future SDTMIG-4.0 standard. One of these "new" domains is LC: Lab Test Results in Conventional Units.
This until now non-official domain was already mentioned in numerous FDA publications, including Technical Conformance Guides (TCGs). From the June 2026 publication:

  

LC was not present in SDTMIG-3.4, so it is clear that the new LC domain has been added on request of the FDA.

The CDISC draft specification of LC states:
"Not all regulatory agencies require the LC domain in submissions. Refer to the specific guidance from the relevant regulatory authority to confirm expectations and timing for inclusion". It does not specifically state that it is "FDA only", but it is suggested by a sentence in the examples: "This example illustrates the US FDA requirement to have both LB and LC domain datasets for the same data within a submission".

What is the LC domain and how does it differ from LB?

If one reads the "assumptions", it becomes clear that an LC dataset is expected to be identical to the LB dataset (exactly the same number of rows, exactly the same number of variables, only with "LB" replaced by "LC" in the variable names), but with all values for "standardized results" using "conventional units" instead of "SI units".

This means that LCSTRESC/LCSTRESN/LCSTRESU must be using "conventional units", whereas LCORRES and LCORRESU must be identical to LBORRES/LBORRESU in the corresponding LB dataset.

Rather surprisingly, the draft specification does not say anything explicitly about LCSTNRLO and LCSTNRHI (reference ranges in standardized units), but from the examples, it shows that these must then also be using the "conventional units".

What are "conventional units"?

So, what are "conventional units"?
A bit of research reveals that there is no formally defined specification of what a "conventional unit" is, nor is there an official list. Usually, "SI units" refer to "molar" units, like "mmol/L", i.e. molar concentrations, whereas "conventional units", refer to "mass concentrations". BUT: there are exceptions … 

For example, FDA regards "g/L" not as a "concentional unit", but requires "g/dL" instead as the "conventional" unit. It seems to consider "g/L" as "SI", although it is not a "molar concentration". This is also visible from the examples in the draft specification.
Another example is "L/L", which is often used for hematocrit measurements. This is regarded as an "SI unit", with "%" as the conventional unit.

As said, unfortunately, there is no formal list of what exactly is a "conventional unit", and what is not. Even the CDISC specification of LC does not provide a definition at all, nor a link to a definition or list.
So, it all seems to be based on "tradition" …

Why this "new" domain?

 The question then arises why CDISC wants to have this "new" LC domain, which is an almost exact copy of LB, with just other values and units for a few variables only.
It is clear that this is then an "FDA-specific" domain, and probably only added by request of the FDA. Did CDISC (once again) give in on the "requirements" of one single regulatory authority, whereas there are much more simple solutions?

The most plausible reason for this new domain, and the FDA requirement to submit a dataset for it according to the FDA-TCG, is that reviewers are not capable to do any unit conversions themselves, or that their software systems are not able to provide such features. So we must do it for them …

The NLM RESTful Web Services

This is a bit surprising, as the NLM (National Library of Medicine), a US institution (!) offers a RESTful Web Service (RWS) to automate such conversions. It can be found at https://ucum.nlm.nih.gov/ucum-service.html.
It uses the
UCUM international standard for unit notation, but for concentrations, CDISC units are for over 95% identical to UCUM. The other way is not true, as CDISC has a "list" of units, whereas UCUM is a "system". So, for concentrations, one can say that the CDISC units are mostly a subset of UCUM.

Let us take an example of the LC draft specification, example 2. This is the only example where LB contains a "molar" concentration, example 1 has no such, though typical, rows:

We find:

  

 Using the NLM RESTful Web Service, this conversion is easy to automate. The RWS request simply is:

https://ucum.nlm.nih.gov/ucum-service/v1/ucumtransform/4.9/from/mmol/L/to/g/dL/MOLWEIGHT/64458

leading to the result:
 

with the molecular weight of human hemoglobin being 64458, which is the molecular weight of the tetramer.
If however a molecular weight of 16000 is used, the result is:  

When we look into example 2 however, we see that the conversion may not been done correctly, depending on which molecular weight was used.

 

From the information in the example itself, it is completely unclear which molecular weight was used for the conversion. We can just guess that 16000 was used. It is however not documented at all.

When one does have the LOINC code available, the RWS request is even easier. For example for glucose in blood, starting from the 5 mmol/L with the LOINC code being 14771-0 and we want to convert to g/dL, the request is:

https://ucum.nlm.nih.gov/ucum-service/v1/ucumtransform/5/from/mmol/L/to/g/dL/LOINC/14771-0

with the result being:

 

Essentially, it looks as FDA reviewers are not capable to do unit conversions in an automated way.
Can it be that they do not have internet access? Or do their tools not support (or do not allow to) use RESTful Web Services? Even within SAS, this is extremely easy to accomplish, like:

 

Remark that I generated this script using AI - I am not a SAS specialist at all ...

Can't FDA reviewers simply implement such a script in order to do the unit conversions themselves? Why do we need to do these for them? It is just another source of possible error ...

That even the CDISC example shows a result that depends on which value for the molecular weight was used (and which is not documented in the file), raises a lot of issues:
For example: how can FDA reviewers check that the value in e.g. LCSTRESN is correct when they can't even use RESTful web services? Will they just believe that everything the sponsor provides in LC is correct? This can easily lead to catastrophe ...

So, what I am expecting, is that once sponsors start really using LC (although I presume they already submit the "unofficial" LC dataset), and FDA reviewers somehow find out that some of the conversions have not been done correctly, they will ask CDISC for yet another domain, this time a non-subject domain, like TC, "Trial Conversion Units", where the sponsor must submit all the conversion factors used in the study to generate LC.

But then, how will or can the FDA reviewers then check that the conversion factors used are correct? Checking this will consume review time ...

In the mean time, patients are dying as it takes too long before their life-saving medication becomes available ...

Other issues with LC

The LC specification states the LCLOINC is a "permissible" variable. However, the FDA requires the LOINC code to be submitted, when available from the lab. So, as this is an "FDA-specific" domain, shouldn't LCLOINC be at least "expected"? Or did the authors of the specification just copy of what is in LB?

The other issue is that when "conventional units" are provided, it can be that the LOINC code also changes. For example for "concentration of glucose in urine", when "molar concentration" is provided, the LOINC code is 15076-3.

When however, "mass concentration" (i.e. "conventional unit") is provided, the LOINC code is 2350-7, with the typical unit being mg/dL. So, if we collected "concentration of glucose in urine" in molar units in LB with LBLOINC=15076-3, and then generate the LC dataset, what should we use for LCLOINC?
The draft specification does not say anything about this …
 

When we look at example 2 in the draft specification (Hemoglobin in blood) which has mmol/L in LB and g/dL in LC, we see that LBLOINC=59260-0 and LCLOINC is also … 59260-0, which is essentially wrong. The correct LOINC code for hemoglobin with units of g/dL is 718-7. So it seems that, just from this one example, it is expected that the LOINC code must be copied from LB to LC, although this leads to an incorrect LOINC code in LC, at least for the "standardized" test.
The draft specification hower does not say anything about this at all.

The specification also says nothing about LCSTNRLO and LCSTNRHI. Must the values also be converted or does not need to copy them from LB? The specification doesn't say it. When however looking at the examples, it becomes more or less clear that they also need to be converted, also this is never explicitly written.

Other minor, but not unimportant issues with the draft SDTMIG-4.0 specification

When studying this draft specification, I noticed a few things for which there were a lot of comments in the first draft publication, but seems not to have been handled at all yet.

- The specification does not explicitly what the "label" for the dataset is. As well for SAS-XPT as for Dataset-JSON submissions the "label" must be provided and also goes into the define.xml under ItemGroupDef/Description.

The specification provides:

 

Ok, we all know from the previous IGs that what is highlighted here in yellow goes into the "label", but it would of course be better if it was stated explicitly - I don't like "interpretation by tradition" …

- The second issue which had a lot of comments was about the designation of "lc.xpt". Essentially, it suggests that the only possible for datasets, whether they are submitted or not, is SAS Transport 5. We all know that it can be expected that the FDA may soon allow submissions in the modern Dataset-JSON format. If "lc.xpt" (and of course others) is kept, it could give the opponents of a modern submission format another argument, like "it is not allowed by the SDTMIG" …
Personally, I think this is a serious issue.


 

No comments:

Post a Comment