Java

JPA and Hibernate Fundamentals

Executive Summary

JPA and Hibernate replace hand-written row mapping with entities. An entity is a plain Java class annotated @Entity with an @Id, mapped to a table by convention or annotation. It is deliberately mutable, with identity semantics, unlike the records that carried rows in the CRUD article. That is because JPA’s contract is objects it can load, change, and track. The EntityManager is the API. For example, persist stores a new entity, find loads by id, and remove deletes. Meanwhile, the persistence context, the EntityManager’s first-level cache, keeps loaded entities managed. So dirty checking flushes your changes automatically at commit, with no update statement written by you.

Bootstrap declares a persistence unit in persistence.xml with the driver, URL, dialect, and, for this article, a generated schema. Then an EntityManagerFactory is created once, and EntityManagers are borrowed per operation like pooled connections. Queries go through JPQL, JPA’s object-oriented query language, so you write b.title instead of books.title. Parameters bind by name, as injection-safe as the PreparedStatement discipline underneath.

Transactions remain yours to manage, and the all-or-nothing article’s rules are unchanged. Finally, some trade-offs earn respect. You must occasionally read the generated SQL, and lazy loading can ambush the unaware. The N+1 query problem issues one query per row, so its fix, JOIN FETCH or batch sizing, belongs to your vocabulary from day one.

Entities: Classes the Persistence Layer Understands

The CRUD article’s Book was a record, so the contrast is the fastest way to understand what an entity is:

import jakarta.persistence.*;

@Entity
public class Book {
    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;                       // database identity: the entity IS this row

    private String isbn;
    private String title;
    private String author;
    private int year;

    protected Book() {}                    // JPA requires it: the provider's door in

    public Book(String isbn, String title, String author, int year) {
        this.isbn = isbn; this.title = title; this.author = author; this.year = year;
    }

    public void renameTo(String newTitle) {   // MUTATION is the point: JPA
        this.title = newTitle;                 // tracks changes to managed objects
    }

    // getters, and equals/hashCode on the id, not every field
}

Read the three deliberate differences from a record. First, the entity is mutable, because JPA’s workflow is load, change, and let the framework write the update. That is the opposite of the records article’s immutable-value discipline. Second, identity is the id field, so equals and hashCode use it. After all, two references to the same row are the same entity regardless of field drift. Third, the protected no-argument constructor exists because the provider constructs instances through reflection. You will meet that implementation detail in every entity you ever write. Fields map to columns by name matching, with annotations for the exceptions. The specification documents the full vocabulary, @Column, @ManyToOne, @OneToMany, when you need them.

The EntityManager and the Persistence Context

The EntityManager is JPA’s unit of work. Your code borrows it per operation, like a connection, and it mediates between your entities and the database:

try (var em = emf.createEntityManager()) {           // borrowed, like a connection

    // CREATE: persist schedules the INSERT
    var book = new Book("978-0134685991", "Effective Java", "Bloch", 2018);
    em.persist(book);                                 // id is assigned, object managed

    // READ: find loads by primary key
    Book found = em.find(Book.class, 1L);             // first checks the context cache!

    // UPDATE: there is no update method. Mutate a MANAGED entity; the flush
    // compares it against its loaded snapshot and writes the UPDATE for you
    found.renameTo("Effective Java, 3rd Edition");

    // DELETE: remove
    em.remove(found);
}

The persistence context is the concept that explains everything surprising that follows. Each EntityManager tracks every entity it has loaded, the managed set, with a snapshot of its original state. At flush time, before the transaction’s commit, the provider compares entities to snapshots and writes exactly the changes. That is why there is no update method: mutation of a managed object is the update. The context is also a cache. A second find of the same id within one EntityManager returns the same object without a query. However, an entity that leaves the context’s scope becomes detached, a plain object whose changes no longer flush anywhere.

persistence.xml and Bootstrap

The persistence unit is the configuration, META-INF/persistence.xml on the classpath, declaring the database and the provider:

<persistence xmlns="https://jakarta.ee/xml/ns/persistence" version="3.0">
  <persistence-unit name="catalog">

    <class>in.imraan.catalog.Book</class>       <!-- listed, or auto-scanned -->

    <properties>
      <property name="jakarta.persistence.jdbc.url"
                value="jdbc:h2:mem:catalog;DB_CLOSE_DELAY=-1"/>
      <property name="jakarta.persistence.jdbc.user" value="sa"/>
      <property name="jakarta.persistence.jdbc.password" value=""/>

      <!-- Hibernate specifics -->
      <property name="hibernate.dialect"
                value="org.hibernate.dialect.H2Dialect"/>
      <property name="hibernate.hbm2ddl.auto" value="drop-and-create"/>
      <!-- ^ schema from entities: examples and tests ONLY. Migrations in prod -->
      <property name="hibernate.show_sql" value="true"/>
      <!-- ^ print generated SQL while learning: your window into the mapping -->
    </properties>
  </persistence-unit>
</persistence>
// bootstrap: the factory once, managers per operation
var emf = Persistence.createEntityManagerFactory("catalog");
try (var em = emf.createEntityManager()) { ... }
emf.close();   // at shutdown

The two flagged properties are the honesty levers. First, hibernate.hbm2ddl.auto generates the schema from entities, which is for examples and tests only. The CRUD article’s migration-tool warning applies with interest. Second, show_sql prints every generated statement, so the mapping stays readable rather than magical, per the Hibernate user guide that documents all of it.

Transactions Still Belong to You

JPA does not manage transaction boundaries for you. The CRUD article’s all-or-nothing rule transfers exactly, and the API for a standalone application is em.getTransaction():

try (var em = emf.createEntityManager()) {
    em.getTransaction().begin();                    // setAutoCommit(false), JPA form

    var book = em.find(Book.class, 1L);
    book.renameTo("Effective Java, 3rd Edition");   // dirty: flush writes the UPDATE

    em.getTransaction().commit();                   // flush, then commit: atomic
}

Without the explicit transaction, persist and mutation flush nowhere or throw TransactionRequiredException, depending on the operation. The honest mental model is the pooling article’s transfer, with the persistence context as the staging area. Entities change in memory, then the flush writes the SQL, and the commit makes it permanent. An exception before the commit leaves nothing behind. The flush-before-commit ordering also explains why reads inside a transaction can see your own unflushed changes. The context is ahead of the database, and the provider keeps the two consistent at the boundaries.

JPQL: Queries Over Entities, Not Tables

find covers primary keys. Everything else goes through JPQL, a query language whose nouns are entities and fields, not tables and columns:

// SELECT over the ENTITY: Book is the class, b its alias, title its FIELD
var query = em.createQuery(
        "SELECT b FROM Book b WHERE b.author = :author ORDER BY b.title", Book.class);
query.setParameter("author", "Bloch");      // named parameters: the PreparedStatement rule
List<Book> books = query.getResultList();

// a projection: selected fields into a record, for read-only views
record BookView(String isbn, String title) {}

var views = em.createQuery("""
        SELECT new in.imraan.catalog.BookView(b.isbn, b.title)
        FROM Book b WHERE b.year >= :year
        """, BookView.class)
        .setParameter("year", 2000)
        .getResultList();

Two properties carry all the value. The named parameters are bound, never concatenated, so the injection discipline is structural here too. Also, the projection form, new Package.View(…) in the SELECT, is JPA’s escape hatch back to the records discipline. For read-only screens, you skip the entity machinery entirely and receive immutable views. It is the mapping layer’s answer to “I just need the data, not the tracking.”

The N+1 Problem: The Trade-Off With Teeth

Every mapping layer writes SQL you did not hand-write, and the most famous way that bites is N+1. Entities relate, such as an Author with a lazy List of Books. The innocent loop queries once for the authors, then once per author for the books. That is one plus N statements where one joined query would have served:

// the trap: looks like Java, costs 51 queries for 50 authors
var authors = em.createQuery("SELECT a FROM Author a", Author.class).getResultList();
for (var a : authors) {
    total += a.getBooks().size();      // each size() may trigger a LAZY load: 1 query each
}

// the fix: FETCH JOIN, one query, everything loaded
var authors = em.createQuery(
        "SELECT a FROM Author a LEFT JOIN FETCH a.books", Author.class).getResultList();

Detection uses the profiling article’s SQL log: show_sql or the pool’s statement counter. The fix is deliberate fetching, with JOIN FETCH for the known paths and batch size configuration for broad cases. The general rule earns its place in every review. An ORM removes the SQL you would have written, but it can silently write worse SQL than you would have. So the generated statement log is part of the code review until the patterns are proven. After all, “it works” is not the same as “it queried once”.

How Real Systems Do This

Hibernate is the default mapping layer of Java services, usually through the JPA specification. In production, JPA and Hibernate follow consistent conventions everywhere. Entities have id-based identity, and transactions sit at the service boundary. Query-heavy screens get read-only views projected into records or DTOs. Also, statement logging in development keeps the generated SQL one glance away. The high-performance outliers skipped JPA for plain JDBC in the JDBC repository shape. They prove the same point from the other side: both choices are respectable when made deliberately for the workload. For detection and production fixes beyond JPA, see N+1 Queries: Detection, ORM Pitfalls, and Production Fixes.

My N+1 story is the standard one, because it is the standard way engineers learn to respect generated SQL. A catalog page rendered fine in development with the seed data: twelve authors, thirteen queries, and nobody counting. Then it fell over in production at real scale. Two thousand authors meant two thousand one lazy loads per page, so the database drowned in single-row lookups. The p99 went from acceptable to embarrassing in one deploy.

The profiling article’s method found it in minutes, because the SQL log showed the same SELECT repeating like a drumbeat. The fix was one JOIN FETCH on the known path plus a team rule. Every new query relationship gets its fetch strategy named in review: explicitly eager, explicitly lazy with a fetch plan, or explicitly projected. The page went back to one query. The rule stayed, and I would hand it to you along with this article. With JPA and Hibernate, the mapping layer writes your SQL now, so reading that SQL is part of writing the code.

Decision Framework

  1. Does the domain have rich relationships and identity-tracked objects? JPA: the mapping, dirty checking, and relationship loading earn their complexity.
  2. Is the work query-shaped, reports, searches, projections? Plain JDBC from the CRUD article, or JPA projections into records: entities would track what you never change.
  3. Are entities being compared or put in sets? equals and hashCode on the id, or the collection misbehaves exactly like the Object article warned.
  4. Is a relationship traversed in the same operation that loaded it? Name the fetch strategy: JOIN FETCH, batch size, or eager, and let review confirm it.
  5. Is a read-only screen the destination? Project into a record view and skip the entity entirely.
  6. Schema from annotations? Tests and examples only: production keeps the CRUD article’s migration discipline.
  7. Generated SQL under suspicion? show_sql on, the profiling article’s method, and the statement count answers before the theories start.

When NOT to Use This

  • Do not use records as entities: JPA needs mutable, identity-based classes with a no-arg constructor, and records are deliberately neither. Keep records for views and messages, exactly as the projection pattern does.
  • Do not enable schema-from-entities in production. The CRUD article’s drift lesson applies, and Hibernate’s generation is not a migration history either.
  • Do not keep EntityManagers alive across requests or operations. The persistence context is a unit of work, whereas a long-lived context is a growing cache of detached expectations.
  • Do not traverse lazy relationships after the context closed. LazyInitializationException is the shape of that mistake, so the fix is fetching inside the work, not holding contexts open.
  • Do not fight the generated SQL blindly. If the mapping keeps producing the wrong shape, the CRUD article’s explicit repository remains a first-class alternative, not a failure.

Common Mistakes

  • The N+1 loop: works with test data, dies at scale, and the statement log is the early warning, per this article’s war story.
  • equals and hashCode over all fields. Two references to one row compare unequal mid-flush, and collections split identities, so id-based identity is the rule.
  • Mutating a detached entity: the change flushes nowhere, silently, and the bug looks like “my update sometimes does not save.”
  • Forgetting the transaction: persist outside a transaction throws or vanishes by provider policy, and the fix is the boundary, not retries.
  • toString or logging over lazy relationships: the string builder triggers a load, sometimes outside the context, and the debugging tool becomes the bug.
  • Caching entities in application fields. A stale detached graph serving as a cache reproduces the concurrent collections course’s races with none of the guarantees.
  • Trusting entity graphs over the wire. Serializing managed entities leaks lazy proxies and schema shape, which is why the projection pattern ships records to the edges.

Key Takeaways

  • JPA is the specification, and Hibernate is the reference implementation. An entity is a mutable, id-identified class the provider can load, track, and write.
  • The EntityManager is the unit of work. Meanwhile, the persistence context tracks managed entities, caches by id, and dirty-checks your changes into generated UPDATEs at flush.
  • There is no update method: mutation of a managed entity is the update, and detach ends the tracking.
  • persistence.xml declares the unit, and show_sql plus a test-only schema keep the generated side visible instead of magical.
  • Transactions remain yours: begin, mutate, commit, and the all-or-nothing rules are unchanged from the pooling article.
  • JPQL queries entities and fields with named parameters, and projections into records are the read-only escape hatch back to immutable views.
  • The N+1 problem is the trade-off with teeth. Read the generated SQL until the fetch strategies are proven, because the mapping layer writes your SQL now.

FAQ

What is JPA in Java?

The Jakarta Persistence API, the standard specification for object-relational mapping in Java: it defines entities, the EntityManager, the persistence context, JPQL, and the persistence unit. Hibernate is its reference implementation, the engine that generates the SQL.

What is the EntityManager in JPA?

The unit of work: borrowed per operation, it persists new entities, finds by id, removes rows, and tracks every loaded entity in its persistence context. Mutating a managed entity and committing writes the update, which is why JPA has no update method.

What is a JPA entity?

A mutable class annotated @Entity with an @Id, mapped to a table, whose identity is the id and whose changes the provider tracks while managed. It is deliberately the opposite of a record: mutable, identity-based, and constructed through a no-arg constructor the provider requires.

What is JPQL?

JPA’s query language over entities and fields rather than tables and columns: SELECT b FROM Book b WHERE b.author = :author, with named parameters bound like PreparedStatement values. Projections, new View(…) in the SELECT, return immutable record views for read-only work.

What is the N+1 problem in JPA?

The performance trap where loading N entities and then traversing a lazy relationship on each issues one query for the list plus one per entity, N+1 statements total. The fix is deliberate fetching, JOIN FETCH or batch configuration, and the statement log is the detection tool.

Conclusion

You now hold JPA and Hibernate in their honest form. The provider tracks entities, and you borrow an EntityManager per operation. Updates are dirty-checked, JPQL binds parameters, and the generated SQL stays visible where it belongs. The trade-offs, mutable entities, lazy loading, and N+1, are now yours to manage deliberately instead of discovering in production.

The next article names a pattern that both this article’s EntityManager work and the CRUD article’s JDBC repository already used: the repository pattern. Its interface hides persistence behind save, find, and delete. As a result, services depend on the contract, and storage becomes an implementation detail you can swap, mock, or test.

Let the mapping layer write the SQL, then read what it wrote. The review line is the whole contract.

Last updated on 12 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *