Planar Chirality
Introduction
Planar chirality is a term used to refer to stereoisomerism resulting from the arrangement of out of plane groups with respect to a plane. A molecule possessing a chiral plane exhibits planar chirality. A chiral plane is not definable as easily as a chiral center or axis. A chiral plane contains as many of the atoms of the molecule as possible. Chirality is due to the fact that at least one ligand (usually more) is not in the chiral plane. Moelcules with planar chirality include ansa compounds, paracyclophanes, metacyclophanes, metallocene and a few trans-cycloalkenes. In the enantiomers, the methylene chain is on either side of the aryl ring or C=C bond. The interconversion between the enantiomers is prevented by the inability of the alicyclic ring, being too small, to swing from one side to the other of the aryl or olefinic plane.
Definition
Chiral molecules have potential plane of symmetry that consist of some atoms. There exist out of plane group perpendicular to the potential plane. One or more substituents destroys the plane of symmetry, perpendicular to the symmetry plane. In the below figure, atoms a, b and x lying in a potential plane, y-z in a plane with atom z orthogonal to the potential plane and substituent Br destroying symmetry plane fall into planar chirality. Implicit in this description is that z is restricted from lying in the plane.
The (R, S) Specification of Planar Chirality
The specification of these compounds is done by application of the following selection rule.
Selection rule: The most preferred atom directly bound to atoms in the chiral plane is selected as the pilot atom (spectator point). It is the first out-of-plane atom linked to the sequence-preferred end of the chiral plane. The sequencing starts with the first in-plane atom directly bound to the pilot atom (the underlined C) and going along the in-plane sequence (marked as a, b, and c) involving the more preferred atom at each branch. The order in which a, b, and c appear when seen from the pilot atom specifies the absolute configuration, i.e., R for clockwise and S for anticlockwise order.
The specification of planar chirality have three steps. First, pilot atom, not in the plane but attached to an atom in the plane on the side of highest priority in the plane e.g.Br. Second, sequence start at the atom in the plane attached to the pilot, follow atoms in the plane towards to the higher priority atoms. Third, view from the pilot atom, if sequence clockwise (R), anticlockwise (S).
Some compounds of planar chirality and their (R, S) specification.

Actually, there are standard CIP nomenclature rules to define the absolute configurations in planar chiral molecules. As described in IUPAC Blue Book 2013. A tetrahedron (tetrahedral stereogenic unit) is derived by connecting the in-plane reference atoms ‘a’ and ‘b’ to the atoms ‘y’ and ‘z’. The chirality sense is determined by conventional rules establishing the priority order ‘a > b > c > d’ or ‘a > b > y > z’ for the cyclophane below. Thus, the descriptor ‘Sp’ is used to denote the configuration and assigned to the carbon atom.
Reflection invariance problems
Since chiral molecules can have two different nonsuperimposable mirror image structures, the 3D-arrangement of groups around a chirality element should have mirror image relationship. The mirror image of R molecule will always be an S molecule. The change of R descriptor to S in the mirror image is called reflection variance (Figure 1) while if R remains R (or S remains S) in the mirror image, that is called reflection invariance. We know that R/S notation are used to assign absolute configuration, so the R form of one enantiomer should become S form in its mirror image. But we can find some special planar chirlaity molecules violate this rule. If two highest ortho substituents are same except configuration on the benzene ring of ansa compounds. When using previous pilot method or standard CIP rule to assign the configuration, reflection invariance is observed.

Ansa compounds
Benzene derivatives having para positions (or meta) bridged by a chain (commonly 10 to 12 atoms long) (Latin ansa, handle). By extension, any arene bridged by a chain constrained to lie over one of the two faces of the arene. Generally, there are three types of ansa compounds, [n.n]cyclophanes, [n]cyclophanes and [n]metacyclophane. "Small" cyclophanes with short chains constitute a model for studies on fundamental aspects of strain and aromaticity. The tension imparted on the whole system generates distortion from aromatic planarity and provokes unusual reactivity behaviors. In addition, the restricted rotation of the aromatic ring may generate planar chirality. The configurational stability of this stereogenic unit is hard to predict and relies upon several factors, such as the length and constitution of the chain and the size of the substituents in the aromatic moiety. In contrast to chiral metallocenes and metal arenes, only one substituent is required to produce planar chirality.

For ansa compounds, the substituents on the benzene ring will affect the planar chirality, if the substituents don't destroy the symmetry plane, they will be achiral. Below figure describe how substituents affect the chirality.

(E)-Cyclooctene
The compound (E)-cyclooctene, which is the first stable cyclic alkene with a trans arrangement of the hydrogen atoms about the double bond. There is a large barrier for the double bond to pass through the ring. Both enantiomers have a two-fold axis of symmetry. Two carbon atoms of the double bond and the directly attached atoms of the substituents (2 carbon and two hydrogen) define the stereogenic plane–the remaining chain can loop around the back from top right to bottom left, or from top left to bottom right. These two forms are enantiomers–remembering that our ultimate criterium for chirality is a mirror image which is not superposable on the original.

Metallocene
Metallocenes are organometallic coordination compounds in which one atom of a transition metal (iron, ruthenium, manganese, osmium, titanium and zirconium) is bonded to the face of two cyclopentadienyl [η5-(C5H5)] (Cp) ligands which lie in parallel planes.
Planar chirality of metallocenes is related to their threedimensional sandwich-like spatial arrangement where a transition metal is situated between two Cp ligands. Cp anions (Cp−) are achiral flat-shaped chemical species which upon disubstitution (R1 ≠ R2) presents two enantiotopic faces. Thus, π-coordination of the planar prochiral Cp− to a CpM+ (metal cyclopentadienyl cation) discriminates the two enantiotopic faces to induce planar chirality in the metallocene complexes. In other words, disubstitution (in 1,2- or 1,3-positions) of one of the Cp ligands of metallocenes generates a chiral plane (coloured in red) where the two sides of the plane are differently occupied.

The nomenclature rule proposed by Schlögl in 1964 to define the absolute configuration of planar chiral ferrocene compounds proceeds in three steps. First, considering the chiral plane, iron is assigned as the pilot atom that is the linked atom being out of the chiral plane and with the highest priority based on the CIP priority rules. Second, the chiral plane is observed from the top, thus positioning the iron atom beneath the plane. Third, the two substituents on the Cp are classified according to CIP priority rules: if the relative orientation from the first-priority substituent to the second one is clockwise, the molecule is assigned as an (R)-stereoisomer; if it is counterclockwise, the molecule is assigned as an (S)-stereoisomer.

The standard CIP nomenclature rules generally used for defining absolute configurations of "central chirality" can be applied to the notation of planar-chiral molecules. For example, one can assume that there are virtual σbonds between the central iron atom and each of the five carbon atoms of the cyclopentadienide core in 1-ethyl-2-methylferrocene. With this assumption, the five carbon atoms of the substituted π-ligand can be regarded as "distorted tetrahedral sp3-carbons" which are centrally stereogenic. Then, based on the standard notation method for stereogenic atoms, the five atoms in the planar-chiral ferrocene can be notated either as (R) or as (S). The enantiomer of 1-ethyl-2-methylferrocene is (1S,2R,3R,4S,5S)-isomer. However, this longsome description can be shortened in most cases. Whereas the absolute configurations of the five stereogenic carbon atoms are synchronized each other, the absolute configurations of only the substituted carbon atoms are written as "(1S,2R)-isomer". Occasionally, the absolute configuration of the highest priority atom based on the CIP rules is mentioned and the other's are omitted as "(1S)-isomer" or "(S)-isomer".
As declared above, an identical stereoisomer can be notated either (R)- or (S)-enantiomer depending on the nomenclature rules used. Schlögl, who proposed the former notation rule in 1964, later recommended the use of the latter nomenclature system for planar-chiral compounds. However, two of the pioneering (and still very influential) papers on planar-chiral ferrocenes (Ugi's in 1970; Hayashi & Kumada's in 1974) employed the former nomenclature method, and thus the former rule has been frequently used in metallocene chemistry. On the other hand, the latter notation rule has been popular in (π-arene)chromium chemistry.
Conclusion
Planar chirality is a special class of chirality, this post introduce the classical type of planar chirality, describe the ways to assign the CIP descriptors, and the reflection invariance problems.
References
- Nomenclature of Organic Chemistry. IUPAC Recommendations and Preferred Names 2013.
- Basic Concepts in Organic Stereochemistry
- Ferrocene derivatives with planar chirality and their enantioseparation by liquid‐phase techniques
- Catalytic asymmetric synthesis of planar-chiral transition-metal complexes
- Planar Chirality
- Planar Chirality: A Mine for Catalysis and Structure Discovery
- Planar chiral [2.2]paracyclophanes: from synthetic curiosity to applications in asymmetric synthesis and materials
- Through a Glass Darkly—Some Thoughts on Symmetry and Chemistry
- The reflection invariance problems in stereochemical nomenclature for absolute configuration
Matched Molecular Pair Analysis
Introduction
In drug discovery, when the most promising lead compound found from screening, it needs to be further improved in one or more properties in the lead optimization process before it can be considered as a clinical candidate. In this scenario, it is about understanding and predicting what effect changing the structure will have on the properties, ideally retaining or enhancing the desirable properties while reducing the undesirable. Matched molecular pair analysis (MMPA) is a promising approach that can be used for this purpose. The term of Molecular Matched Pair Analysis was coined by Kenny and Sadowski in 2004 for a special case of QSAR, and now it is widely used in drug design processes.
What Is Matched Molecular Pair Analysis?
One definition of MMPA can be described as "identifying every pair of molecules that differ only by a particular, well-defined, structural transformation in a database of measured properties and computing the corresponding change in property". Such pairs of compounds are known as matched molecular pairs (MMP). Because the structural difference between the two molecules is small, any experimentally observed change in a physical or biological property between the matched molecular pair can more easily be interpreted. MMPA inspire people to think about medicinal chemistry differently, can help chemists to discover new effects, provide insights, understand the relations between structures and properties.
Application
The assumption that the effect of chemical substitution can be generalized, is inherently assumed in all QSAR methods, including the MMP approach, successfully highlighted by the work of Lipinski et al. who correlated physicochemical properties to oral bioavailability. With the increasing availability of public databases containing millions of structure–activity-relationship (SAR) or SPR data, multiple papers have been published applying MMP concept to: ADME, bioisosterism, aqueous solubility, plasma protein binding, oral exposure, logD, potency, intrinsic clearance, herG and P450 metabolism, in vitro UGT (Uridine 5′-diphosphoglucuronosyltransferase) glucuronidation clearance, half-life, selectivity against off-targets, impact of N- and O-methylation on aqueous solubility and lipophilicity or mode of action; the analysis differing only in the MMP algorithm used.
Activity Cliffs
One interesting subset of matched molecular pairs is those in which the change in property is surprisingly large. A number of researchers, most prominently the Bajorath group in Bonn, have reported the insight that these pairs, which they call activity cliffs, can bring. Activity cliffs are generally defined as pairs of structurally similar compounds having large differences in potency. Notable recent contributions include using Hussain and Rea’s fragment and index method to identify matched series within all the high-confidence ChEMBL KI data relating to activity against human targets. These reveal that coordinated activity cliffs can be identified in which the same structural change causes the same large change in property across several chemical series. A similar method can be used to explore activity within large sets of screening hits. Using MMPA to study the activity cliffs can discover new insights and provide an alternative source for further chemical exploration. In contrast to traditional SAR analysis, where similar compounds are assumed to have similar properties, activity cliffs describe the substitution pattern with the most impact upon a small structural change.
For example, above is a representative MMP-cliff, "6600nM | 6" indicates that the potency value of the compound is 6600 nM and the number of non-hydrogen atoms comprising the differentiating fragment is 6.
MMPA Algorithms
Generally, there are two broad classes of methods for identifying matched molecular pairs, supervised and unsupervised algorithms. These depend upon whether the pairs are found using manual intervention or automatically.
Supervised
In supervised methods the chemical transformation that generates the MMP is predefined, the SMARTS and SMIRKS nomenclatures are used to define the chemical transformation widely, SMARTS strings have the great advantage that they can be encoded to be extremely specific or general but provide exquisite control in terms of chemical structure. The advantage with supervised methods lies within the precise control of the definition of the MMP to address a particular question. On the other hand, these methods cannot find new and surprising MMPs in the way that unsupervised methods can. In the first publication of MMPA, Kenny and Sadowski described how a molecular editor could be used to identify matched molecular pairs in which substituents are added to a benzene-like ring and analyze the effect of substituent on aqueous solubility.
Unsupervised
Supervised methods to find matched pairs have many limitations, it isn't user friendly and need amount human time to encode the fragment pattern and can't find new MMPs not in the scope of predefined patterns. So a new approach was needed in which matched pairs could be identified automatically. Therefore, a number of unsupervised approaches have also been developed. And they can be divided into two types: fragment and index or maximum common subgraph (MCS).
Fragment and Index
Hussain-Rea fragmentation and index algorithm is the first efficient, unsupervised MMPA algorithm. This algorithm mainly includes two steps:
- Fragmenting. Fragment each molecule in the data set by all possible single, double, and triple cuts.
- Indexing. Generate a key-value store with fragments as keys and values as cores.
In the first step, each molecule is fragmented, by breaking selected bonds. Hussain and Rea achieve this by defining the bonds to be broken using a SMARTS pattern. They aim to have this pattern be specific to acyclic single bonds. The SMARTS pattern suggested is "[*]!@!=[*]" that signifies an atom type joined to any atom type by neither a bond that is in a ring nor a double bond. If a molecule has multiple possible bonds to cut, bonds are broken one at a time, and the resulting fragments are then stored as canonical representations (for example SMILES strings or hash code).
In the above case, a single, acyclic bond within a connected molecular graph (bromobenzene, left) is removed, yielding two fragments (right, phenyl and bromo). The benzene ring can be regarded as a substituent of bromo. Likewise, bromo can be regarded as a substituent of benzene. This symmetry plays a prominent role in the Hussain-Rea algorithm.
Breaking two bonds generate three fragments, one of which necessarily connects to both of its mates when the cuts are reversed. This central fragment is always designated as the core.
Triple cuts will yield four fragments. Here there are two possible results: (1) the third cut is made inside the core; or (2) the third cut occurs outside the core. Result (2) isn't considered in this situation, because triple cuts yields two cores are not allowed by the Hussain-Rea algorithm.
Quadruple (or more) cuts in the fragmentation are possible. However, in the author's opinion, it is only likely to result in a small increase in the number of unique MMPs found.
The next stage of the algorithm is to index these fragments. For single cut, the two resulting fragments formed (i.e., fragment X and fragment Y) are both canonicalized and added to the index. First, fragment X is added as a "key" into the index with fragment Y as its "value". The converse is also carried out; fragment Y is added as a key with fragment X as its value. An identifier for the compound is also stored in the value of the index, so the value can be a tuple (core, ID). In the double cut scenario, the fragmentations result in a core and two terminal fragments. The core is stored as the value with the dot-disconnected unique identifier (in the original publication, using the canonicalized SMILES) of the terminal fragments as their key. Similarly, in the triple cut example, only fragmentations that result in a core and three terminal groups are stored in the index.
After generating the indexes for all molecule fragments, the MMPs could be found directly. In each Key-values map, for each pairing of members in values, generate a matched pair.
The SMARTS definition that identifies which bonds are to be fragmented is not ideal in the original publication and can lead to fragmentations and grouping into sets of pairs that chemists would not normally consider to be chemically sensible. For instance, fragmentation of the single bond in amides, esters, or sulfonamides and acyclic triple bonds would happen if the original SMARTS proposed by Hussain and Rea was to be applied. Consequently, a refined SMARTS definition has been proposed by Wirth and coworkers and is "[#6+0;!$(*=,#[!#6&!R])]!@!#!=[*]". They also proposed a second rule that allows fragmentations of exocyclic double bonds: "[#6+0;R]=[#7&!R,#8&!R,#16&!R,#6+0&!R]". This definition signifies any atom type connected via a bond that is neither in a ring nor double nor triple to a carbon atom that is not charged and that in turn is not double- or triple-bonded to an atom of any element other than carbon (unless that atom is in a ring).
Maximum Common Subgraph
Another class of MMP algorithms is based the maximum common subgraph (MCS),which perform a pairwise comparison within a compound data set to find the MCS between each pair of compounds and define the fixed part. The atoms within a pair of compounds that are not part of the MCS are then analyzed to determine if they constitute a single-point change. This class of MMP identification algorithms is capable of potentially finding all MMPs within a compound data set, but because of the computation expense of MCS algorithms coupled with the O(n2) nature of the MMP identification (because of the pairwise comparisons that need to be performed), the algorithms are computationally expensive. Therefore, to alleviate this limitation, heuristics are performed between pairs of compounds before the MCS or the MMP identification step is carried out (clustering and topological similarity), which may result in certain MMPs not being found. The high computational cost of running these MMP identification algorithms means they are very difficult to apply on large compound data sets.
MCS MMPA algorithm
Conclusion
Matched molecular pair analysis is a powerful tool in drug discovery and other domains. The methods used to identify matched pairs range from manual inspection, through supervised methods to unsupervised methods. The MMP framework allows one to study numerous properties (most commonly binding affinity or potency) and to rationalize the design of the next compound to make within a series.
References
- Structure Modification in Chemical Databases
- Computationally Efficient Algorithm to Identify Matched Molecular Pairs (MMPs) in Large Data Sets
- Matched Molecular Pair Analysis
- Matched Molecular Pairs
- Matched Molecular Pair Analysis in Short: Algorithms, Applications and Limitations
- MMP-cliffs: systematic identification of activity cliffs on the basis of matched molecular pairs