.
Showing posts with label attribute. Show all posts
Showing posts with label attribute. Show all posts

Tuesday, October 31, 2017

Eclipselink 2.5.2 (JPA 2.1.0): Entity Cloning

This discussion made use of JPA 2.1.0 and Eclipselink 2.5.2. Specs and implementation and may change in future releases. It might be helpful to test with different versions by (locally) modifying the pom.xml file that appears in the github repository, Eclipselink Entity Copying, that contains in-depth code and discussion for the concepts relevant to this post.

Also, here is another great reference for when reading through the code: Eclipselink Attribute Groups

Among other cases, the need to clone entities JPA appears under the following conditions:
  • The absence of DTOs (Data Transfer Objects) - which may arguably represent the lack of a clean separation between the model layer and the controller layer
  • The need to protect entities from being polluted by changes, especially in multipart transactions
  • The need to simply persist or manage duplicate information

Although it can be argued that using a JPA provider to this extent is yet another antipattern, we won't talk about that kind of stuff here.

Anyway, say that the need arises for us to clone our entities. Of course, the easiest (but not necessarily the most succint) thing to do is to implement the Cloneable interface and code away on how we are to handle what.

Suddenly, we are hit with some realizations:
  • What if our entities traverse deeply?
  • What if our entities are only partially fetched?
  • How do we want to handle unfetched attributes?

This is when we find out how DTOs fall short. If JPA is to still be used in querying for the required information for, say, a view, then DTOs have to be smart to some degree. It might decided that different views that use information from the same entity would warrant per-view DTOs. DTOs might also have to be made aware of which attributes it should copy from the entity, especially when certain fetch optimizations (attribute narrowing via Eclipselink FetchGroups or JPA FetchGraphs) were used, since violating such optimizations usually lead to the loss of the advantage of their use in the first place.

Luckily, we're using Eclipselink, which has a nifty tool that we can use when in such a predicament: CopyGroups (and the JPAEntityManager.copy method). Honestly though, it won't be the most sturdy of swiss army knives, but the tool we'd be using a lot is luckily also the sharpest one in the set.

Eclipselink CopyGroups and JPAEntityManager.copy()


The main entry point to using this feature is in the following method of the org.eclipse.persistence.jpa.JpaEntityManager class:
It returns Object from a method without generics, so we still have to cast it. Also, the entityOrEntities parameter can also accept a Collection of the same type of entity. Also notice that the method accepts an AttributeGroup; it internally transforms this into a CopyGroup if it not already one. We'll usually pass CopyGroups when we use this method anyway.

We can obtain an instance of a JPAEntityManager via two ways:
  • Directly casting an EntityManager instance (make sure it runs on Eclipselink)
  • Calling unwrap(JpaEntityManager.class) on an EntityManager (again, it should run on Eclipselink)

With that, doing the actual copying is pretty much covered.

What we have to actually be familiar (and careful) with are the configuration options for the CopyGroup we pass to the copy method.

Experiment-discussion


The meat of this discussion actually appears in the test class found in this github repository:

Eclipselink Entity Copying

Simply run the "mvn test" Maven command from the directory that contains the pom.xml file to see if all the test pass (they should). After this, read the code found in the only test class under src/test/main/....

For this post, I'll just leave a summary of the discussion in the code for our reference.

Summary: CopyGroup Configuration


A CopyGroup has two main points of configuration:
  • cascade level, which defines which types (and not numerical depth) of associations the copying should include
  • declared attributes (as it is an AttributeGroup) which it should consider when cloning (only considered when the cascade level is Cascade Tree)

General Considerations

  • Whenever an attribute is added to a CopyGroup, its cascade level is set to CascadeTree. As CascadeTree is only depth that considers attributes, be mindful when adding attributes to a CopyGroup
  • When a FetchGroup, a type of AttributeGroup used for query optimization, is turned into a CopyGroup (via the toCopyGroup() method), it is automatically set to CascadeTree
  • Primary key and version columns can optionally be omitted from the copies via the CopyGroup configuraion methods setShouldResetVersion(boolean) and setShouldResetPrimaryKey(boolean). These options behave differently, according to the cascade level configured
  • Copies can back-reference; that is, when copying with circular references, same entities with the same key share the same reference

Cascade Level Options

  • Cascade All Parts
    • set via the CopyGroup method cascadeAllParts()
    • does not consider attributes it contains
    • copies ALL associations; initializes them if need be
      • For entities and associations that have been or are to be partially fetched, their respective copies would only have copied the attributes for partial fetching (as probably declared via FetchGroup when querying for the original)
      • For associations that were not declared in such a partial fetching scheme, they would be initialized as default
      • (i.e.) ALL associations would still be initialized; only, those that were declared to have a FetchGroup might lack some BASIC attributes
    • when an unfetched BASIC attribute is encountered, its corresponding value in the copy will be null
    • initialization triggered by copying affects the original; i.e. if an association was initialized via a query triggered by copying, then it also becomes initialized in the original
    • because ALL associations are initialized, it might not be worth using this cascade level for heavily associated entities
    • if the group is configured with setShouldResetPrimaryKey(true), the keys will only be reset if none of them are associations (all or nothing)
  • Cascade Tree
    • set via the CopyGroup method cascadeTree()
    • the cascade level is automatically set to Cascade Tree when an attribute is added to the CopyGroup
    • when a CopyGroup is obtained via a toCopyGroup() on a FetchGroup, the resulting CopyGroup uses the Cascade Tree level
    • when passing a CopyGroup without attributes, copying will involve "all attributes", though when it comes to associations, it is still unpredictable (needs more testing); it won't be probable that an empty CopyGroup will be used with the CascadeTree level anyway
    • when accessing an attribute/association from the copy that is not declared in the CopyGroup (which is not empty), an IllegalStateException is thrown;
      • this can be useful for adjusting/optimizing FetchGroups (use them as CopyGroups)
      • be careful when turning FetchGroups into CopyGroups:
        • if a complete CopyGroup is desired, then take it from the FetchGroup that is manually configured and passed as a query hint
        • FetchGroups taken from resulting entities (after being casted to FetchGroupTracker) have been broken down so that they only describe the entity it was taken from
    • if copying triggers initialization queries, then the original entities are affected as well
  • Cascade Private Parts
    • set via the CopyGroup method cascadePrivateParts()
    • supposed to behave like Cascade All Parts, except it cascades only associations annotated with @org.eclipse.persistence.annotations.PrivateOwned
    • it worked the other way around in the tests - all but the PrivateOwned association was cascaded
    • still unpredictable; needs more testing
  • Cascade None
    • set via the CopyGroup method cascadeNone()
    • still initialized associations, thus straying from its name and contract
    • still unpredictable; needs more testing

Unfortunately, only Cascade All Parts and Cascade Tree can be help up to their intention to a usable degree - luckily, Cascade Tree is the level that would see the most use.

In the end, the only CopyGroups actually worth using (fortunately, it should also be the common use case) are those derived from manually configured FetchGroups, or cascade-tree-level groups that were manually built with careful consideration.

Admittedly, this time, it probably seems like a disappointing turnout - one where we are only given a limited number of options.

Perhaps with this we can help each other dig deeper into this feature and learn more about it, or even have the guys over at Eclipselink help us with it.

In any case, once again, hope this helped. Thanks!

Sunday, September 24, 2017

Eclipselink 2.5.2 (JPA 2.1.0): Determining Fetch State

This discussion made use of JPA 2.1.0 and Eclipselink 2.5.2. Specs and implementation and may change in future releases. It might be helpful to test with different versions by (locally) modifying the pom.xml file that appears in the github repository Eclipselink Fetch State Experiment that contains in-depth code and discussion for the concepts relevant to this post.

Also, here is another great reference for when reading through the code: Eclipselink JPA 2.0 Persistence Utils

This "issues" disucussed here do not seem to have been resolved yet as of Eclipselink 2.6.x releases.

Lazy-loading (in various degrees) definitely benefits optimization, and being able to tell whether certain attributes of JPA entities are loaded or not pretty much lies in cusps between such decisions.

At the forefront, JPA actually provides a way to inspect an entity (or its attributes) to determine its fetch state - whether they are loaded or not - through PersistenceUtil.

It can be invoked via the following code:
Looks simple enough. However, is it reliable?

Because it relies on provider implementation, that greatly depends.

To be fair, aside from the method semantics, JPA (as of 2.0, and even in 2.1) provides specification described in the comments in an interface used in the implementation internals of PersistenceUtil, ProviderUtil:
  • isLoadedWithoutReference and isLoadedWithReference, both with arguments (Object entity, String attributeName)

    • "If the provider determines that the entity has been provided by itself and that the state of the specified attribute has been loaded, this method returns LoadState.LOADED."

    • "If the provider determines that the entity has been provided by itself and that either entity attributes with FetchType.EAGER have not been loaded or that the state of the specified attribute has not been loaded, this methods returns LoadState.NOT_LOADED"; and

    • "If a provider cannot determine the load state, this method returns LoadState.UNKNOWN."

    • These two methods are differentiated in that WithReference is permitted to obtain/initialize a reference, whereas the other is not. Note that Eclipselink does not obtain a reference for either one anyway.

  • isLoaded(Object entity)

    • "If the provider determines that the entity has been provided by itself and that the state of all attributes for which FetchType.EAGER has been specified have been loaded, this method returns LoadState.LOADED."

    • "If the provider determines that the entity has been provided by itself and that not all attributes with FetchType.EAGER have been loaded, this method returns LoadState.NOT_LOADED"; and

    • "If the provider cannot determine if the entity has been provided by itself, this method returns LoadState.UNKNOWN."

    • This method is also not permitted to obtain/initialize references.

To put it simply, JPA specifies that an entity is loaded if all the attributes and associations configured to be EAGER by "DEFAULT" have been initialized; if the entity is found to be loaded, then checking for attributes can be done properly and predictably.

The word "default" is stressed here because it actually describes explicit configuration via ORM XML or annotations - pretty much whatever you define at the beginning.

Watch how your provider implements the specification metioned in the comments in ProdivderUtil.

Eclipselink holds true to this (tested in version 2.5.2), and only to this extent. When dynamic configuration is done through runtime application of FetchGroups, PersistenceUtil becomes completely unusable.

Experiment-discussion


The meat of this discussion actually appears in the test class found in this github repository:

Eclipselink Fetch State Experiment

Simply run the "mvn test" Maven command from the directory that contains the pom.xml file to see if all the test pass (they should). After this, read the code found in the only test class under src/test/main/....

With all the technicalities discussed in the github repository linked previously, I'll just leave y'all with a summary of what was in there.

Summary


For Eclipselink, there are actually two main ways to accurately determine entity/attribute/association fetch state. Furthermore, they only work for entities that have been woven. Then again, weaving is what enables lazy loading, and we'd only need to determine fetch state at all if lazy loading was enabled. Anyway, here they are:
  • org.eclipse.persistence.queries.FetchGroupTracker, the simpler (but not the best) way
    • FetchGroupTracker is one of the interfaces that entity weaving adds to your entities. It tracks the FetchGroup used when the specific entity was loaded. A FetchGroup is simply a group of attributes used by Eclipselink to specify which attributes and associations a query or fetch should use.

    • FetchGroupTracker has the _persistence_isAttributeFetched(String attributeName) with which the load state of an attribute can be determined.

      Do this by simply casting the entity to FetchGroupTracker, and use the method accordingly:
    • Now, the problem (or maybe the cool thing) here is that this method is the counterpart of PersistenceUtil; it is unreliable when a FetchGroup is not present for the entity - this means that the entity was fetched using default configuration, where no basic attributes were made LAZY (if a basic attribute was made LAZY, it would have had a default FetchGroup). So if PersistenceUtil works on entities that used defaults while FetchGroupTracker works on entities that used a custom FetchGroup, perhaps they can be made to work together to cover each other's weaknesses (this is only an option; the better ways are described below).
  • org.eclipse.persistence.internal.jpa.EntityManagerFactoryImpl, the actual correct way
    • Even though Eclipselink's PersistenceUtil uses the relevant EntityManagerFactoryImpl methods internally, PersistenceUtil fails due to some of the logic written to follow JPA's specification. However, using EntityManagerFactoryImpl's various isLoaded(...) methods actually work properly (isLoaded(entity) follows JPA's definition of a loaded entity; you'll be using the more attribute-specific overloads).

    • The following snippet describes the methods in question:


And that's pretty much it. Hope this helped!